Memory and task-state control
Several papers give vision-language-action (VLA) policies explicit control over task progress. Harness VLA treats a frozen VLA as a short, retryable contact skill, while a planner handles grounding, transport, staging, and failure recovery. It reaches 82.4% on LIBERO-Pro, compared with 50.0% for the direct frozen baseline. TFP stores an episode-local belief and updates it using elapsed time and interaction events; real-robot object-swap success rises from 3/20 to 15/20. LEEVLA adds task-relevant region weighting and latent future-feature prediction during training, reaching 98.2% on LIBERO without extra inference cost.