Dexterous VLA policies
RLDX-1 treats dexterous manipulation as a multi-signal control problem. The policy adds video motion, a small memory store, and tactile or torque inputs to a Qwen3-VL-based VLA. Its Multi-Stream Action Transformer keeps cognition, proprioception, and physics streams separate before cross-stream attention combines them for action prediction.
The reported gains are largest on tasks where current-image policies struggle. RLDX-1 reports 86.8% success on ALLEX humanoid tasks, while π₀.₅ and GR00T N1.6 are around 40%. On ALLEX Object-in-Box Selection, a memory-heavy task, it reports 91.7% success while the two baselines are in the 30% range. The report also ties deployment to latency: inference optimization cuts per-step latency on an RTX 5090 from 71.2 ms to 43.7 ms.