Cross-embodiment data and adaptive sensing
Scale claims are tied to alignment across robot bodies and sensor choices. Qwen-RobotManip maps different robots into a common state-action template, uses binary masks for missing dimensions, and predicts camera-frame end-effector deltas. Its reported corpus is about 38,100 hours, including synthesized human-to-robot data across 15 bimanual robot configurations, with real-robot validation on AgileX ALOHA, Franka, UR, and ARX.
MuseVLA addresses a different generalization gap. It chooses thermal, acoustic, mmWave, or RGB sensing based on the instruction and scene, then turns the selected signal into a grounded sensor image. On real dexterous-hand tasks, the version with synthesized pretraining reports 80.6% average success on seen tasks and 66.7% on unseen tasks.