Failure recovery and long-horizon memory
Reliability is defined through what the policy does after the task starts to go wrong. RePO-VLA trains on success, failure, and recovery rollouts with separate labels, then reports average adversarial success rising from 20% to 75% and up to 80% in scaled real-world trials. The key detail is supervision over adverse states, contact drift, and useful failure prefixes, not only clean demonstrations.
ECHO attacks the same long-horizon problem through memory. It stores successful subgoal segments in a hierarchical memory and retrieves them during inference. On LIBERO-Long, it reports 93.5% success versus 80.7% for the vanilla π0 baseline. Together, these papers make recovery and memory measurable parts of VLA execution quality.