Video models are becoming robot planners
Video generation is showing up as a planning module for manipulation, not just a data source. Veo-Act uses Veo-3 to predict a future motion sequence, then hands control to a low-level vision-language-action policy during contact-heavy interaction. The reported gains are large on ambiguous scenes and dexterous execution: average success rises from 45% to 80% across the tested sim and real settings, and real-world pass-by interaction improves from 2/13 to 11/13. The paper also makes the limit clear. Video prediction alone can sketch the task, but precise control still needs a reactive action policy.