World-action models for manipulation
Vision-Language-Action (VLA) policies are being built around predicted 3D structure and future contact. STARRY jointly denoises future spatial-temporal latents and actions, then uses predicted depth and end-effector geometry to bias attention toward handles, openings, contact surfaces, and nearby obstacles. It reports 93.82% clean and 93.30% randomized success on 50 RoboTwin 2.0 bimanual tasks, with real ARX R5 experiments averaging 70.8% success across three tasks.
X-WAM takes a broader world-action route. It predicts multi-view RGB-D video, robot states, and 32 future actions in one diffusion model. Its asynchronous denoising schedule lets actions decode with fewer steps than video, which matters for closed-loop control. The paper reports 79.2% average success on RoboCasa and 90.7% on RoboTwin 2.0 Randomized, while also evaluating visual and geometric prediction metrics.