Cross-embodiment motion pretraining
One paper targets a data bottleneck in generalist VLA training: action labels are scarce for robots, while human egocentric manipulation video is abundant. The method learns masked latent action tokens with a disentangled VQ-VAE, using physical masks to separate foreground motion from scene background. A Prismatic-7B vision-language model then predicts those tokens before robot adaptation.
The reported gains are concrete. On LIBERO, the full method reaches 91.8% average success, ahead of OpenVLA at 76.5% and Diffusion Policy at 72.4%. On RoboTwin 2.0 dual-arm simulation, it reaches 67.7% average success across 10 tasks. The downstream setting uses about 50 trajectories per task, so the claim is about cheaper adaptation after unlabeled video pretraining.