Compact spatial foresight for VLA manipulation
ConsisVLA-4D treats spatial consistency as an inference budget problem. It keeps 32 instruction-relevant visual tokens, aligns them across multiple camera views, and stores geometry in compact latent tokens. The reported gains are tied to both accuracy and speed: 21.6% better performance and 2.3× faster inference than OpenVLA on LIBERO, plus 41.5% better performance and 2.4× faster inference on real robot platforms.
T³VF addresses a different failure point in visual-foresight VLA models. It compares a predicted future image with the later observed image, then updates only learnable query tokens when action variance is low. On LIBERO-Plus with perturbed training, it raises Mantis average success from 49.3% to 52.1%, with larger gains on camera and lighting perturbations.