Predictive and long-context control
Prediction is becoming part of the policy state rather than a separate planning output. RoboTTT compresses up to 8K timesteps into fast weights without latency growing with context length. It reports 79% average completion on three bimanual assembly tasks, versus 42% for its single-step baseline. FoMoVLA instead predicts a future feature state and sparse point trajectories; it reaches 97.6% on LIBERO-Long with 9.4 ms median overhead. Lumo-2 provides a third formulation by aligning latent world dynamics with action, vision, and language, though its available evidence does not include numerical task margins.