Cross-embodiment VLA training
Generalist robot policies are being trained to share high-level manipulation concepts while still producing robot-specific controls. ZR-0 is the clearest example. It uses dense embodied chain-of-thought labels during training for scene description, task progress, future plan, target boxes, and discrete action tokens. At inference, it skips text generation and outputs continuous action chunks through a diffusion action expert.
The measured claim is concrete. ZR-0 reports 97.8% average success on LIBERO, with ProcCorpus-60M covering about 60 million frames, 1,000 hours, and more than 400,000 trajectories. The same daily trend also groups this with reward-free test-time improvement and trajectory memory, showing that policy scale is being tied to executable action timing and state history.