Long-context and future-aware control
RoboTTT compresses up to 8K visuomotor timesteps into fast weights, keeping latency constant as history grows. It reached 79% average completion on three real-robot assembly tasks, versus 42% for its single-step baseline. FoMoVLA takes a complementary route: it trains compact future-state tokens together with sparse point trajectories, then removes most auxiliary machinery at inference. It reached 98.8% average success on LIBERO with 9.4 ms median overhead. Together, these studies make temporal information operational through compact internal state rather than an ever-growing observation buffer.