World models inside robot policies
Several papers use future-state prediction as a training signal or as part of action choice. WLA predicts a textual subtask, a compact physical transition, and an action chunk in one policy. Its world-modeling loss has a measurable effect: removing it lowers RoboTwin Clean success from 92.94% to 90.98% and LIBERO average success from 98.6% to 97.9%.
The same idea appears beyond tabletop manipulation. WorldFly couples future video latents with navigation actions for low-altitude UAV control. On its Urban Canyon Traversal benchmark, it reports 87% success on seen intersections and 31% on harder unseen intersections, beating OpenFly and Pi-0-UAV on the reported metrics. DexFuture uses predicted future hand-tool-object targets for bimanual tool use and runs at 60 Hz, avoiding slow online planning over high-dimensional hand actions.