RLDX-1 is a general-purpose Robot Foundation Model designed for dexterous
manipulation. Powered by a Multi-Stream Action Transformer (MSAT), it
seamlessly unifies multimodal perception (visual + tactile), high-DoF
actuation, and memory-aware decision-making in a single architecture. RLDX-1
achieves state-of-the-art performance across diverse simulation benchmarks
and is fully validated on real-world hardware.
This repository hosts RLDX-1-PT — a foundation checkpoint pretrained on
a broad mixture of public manipulation corpora, from which all downstream
RLDX-1-{FT,MT}-* releases finetune. Use it as your starting point for new
embodiments and tasks.
RLDX-1 architecture
Highlights
Multi-Stream Action Transformer (MSAT). Cognition, physics, and
action each get a dedicated stream coupled by joint self-attention —
an extension of MM-DiT to action modeling.
Motion awareness. Multi-frame observations + a motion module
capture temporal dynamics; intermediate VLM layers compress video
tokens to keep the policy efficient.
Long-term memory. A memory module fuses past cognition features
with the current ones for history-grounded decisions beyond a short
multi-frame window.
Physical sensing. Tactile and torque enter as a dedicated physics
stream; the decoder is jointly trained to predict future physical
signals.
Three-stage training. Pre-training (generalization) → mid-training
(functionality) → post-training (task adaptation), with synthetic data
augmenting rare manipulation scenarios.
Real-time inference. Static graph capture + custom fused kernels
bring the all-modality model to 43.7 ms / step on RTX 5090
(1.63× speedup, >22 Hz).
Released Checkpoints
This card describes RLDX-1-PT (foundation). The full RLDX-1 model family:
RLDX-1-PT is pretrained on a multi-source mixture, so for direct inference
pair it with the embodiment tag matching your data source — e.g.
OXE_FRACTAL, OXE_BRIDGE_ORIG, OXE_DROID, GALAXEA, AGIBOT_GRIPPER,
AGIBOT_DEXHAND, NEURAL_GR1, HUMANOID_EVERYDAY_G1,
HUMANOID_EVERYDAY_H1, etc. For custom robots, finetune.
Pretraining data: A mixture of public manipulation corpora, covering
27 Open X-Embodiment (OXE)
datasets (DROID, Bridge, Fractal, Language Table, …) plus
Galaxea, AgiBot World
(Gripper + Dexhand), ActionNet, Neural-Curated GR-1 humanoid trajectories,
and Unitree G1 / H1 from
HumanoidEveryday.
Intended use. Research on robotic manipulation, finetuning on custom
embodiments, simulation benchmarking, and non-commercial real-robot
deployment under the conditions of the RLWRLD Model License v1.0.
Out of scope. Commercial deployment, military or weapons applications,
non-consensual surveillance, and any use that violates applicable laws or
regulations. See LICENSE.md §3.5 for the full list.
Limitations. Performance depends heavily on embodiment match and data
distribution. The pretrained checkpoint is OXE-conditioned and is not
guaranteed to work zero-shot on novel embodiments without finetuning.
Memory, motion, and physics modules are dormant in RLDX-1-PT and only
activate when the corresponding flags are wired during finetuning (see
RLDX-1-MT-ALLEX).
Citation
bibtex
1@article{rldx2026,
2 title={RLDX-1 Technical Report},
3 author={Kim, Dongyoung and Jang, Huiwon and Koo, Myungkyu and Jang, Suhyeok and Kim, Taeyoung and others},
4 year={2026},
5 note={RLWRLD},
6 eprint={2605.03269},
7 archivePrefix={arXiv},
8 url={https://arxiv.org/abs/2605.03269}
9}
License
Released under the RLWRLD Model License v1.0 — a non-commercial license
with attribution and share-alike requirements. See LICENSE.md for
the full text. By using this model you agree to those terms, including the
use restrictions in §3.5.