Latent intent becomes the main VLA design point
DIAL treats future visual features as the control interface. The vision-language model (VLM) predicts a latent future at horizon 16, and a separate policy turns that forecast plus the current observation into a 16-step action chunk. The concrete payoff is data efficiency: the paper reports state-of-the-art results on RoboCasa GR1 Tabletop with 2,400 trajectories where prior full-data runs used 24,000. It also extends the recipe beyond one robot setup, using 27,419 EgoDex human trajectories for zero-shot generalization tests and reporting real-world transfer on the IRON-R01-1.11 humanoid.