Cross-embodiment scale and sensor choice
Qwen-RobotManip treats robot diversity as an alignment problem. It maps different arms, grippers, cameras, and action spaces into a shared state-action template, then trains on about 38,100 hours of manipulation data. Most of that scale comes from a human-to-robot synthesis pipeline that renders egocentric human demonstrations into 15 bimanual robot configurations. The report claims first place on RoboChallenge Table30-v1 generalist track and reports real-robot validation across AgileX ALOHA, Franka, UR, and ARX platforms.
MuseVLA adds a different kind of generalization pressure: the policy must decide when RGB is insufficient. It selects thermal, acoustic, mmWave, or no extra sensor from the instruction and scene, converts the chosen measurement into a grounded sensor image, and feeds it back into the VLA. On real dexterous-hand tasks, synthesized pretraining reaches 80.6% average success on seen sensor-guided tasks and 66.7% on unseen tasks.