World-action models for deployable robot control
World-action models are being evaluated as control systems, not only as predictors. MotuBrain joins future video latents and action tokens in one diffusion model, then cuts reported end-to-end latency from 4.90 seconds to 0.09 seconds. Its RoboTwin 2.0 scores stay above 95% in both clean and randomized settings, which makes the latency result central to the claim.
Being-H0.7 takes a lighter route. It trains latent tokens with future observations, then removes the future-aware branch at inference. The policy keeps the action-oriented latent state and avoids test-time video rollout, with reported deployment at 3–4 ms per step. The survey paper gives the common definition behind these systems: a robot world model predicts future states or observations conditioned on current state, actions, and optional language.