VLA failure recovery and long-horizon control
Two papers attack execution drift directly. RePO-VLA trains on success, failure, and recovery rollouts with separate labels, then deploys with a fixed high value condition and no online failure detector. Its reported adversarial success rises from 20% to 75% on average, with FRBench-Sim covering 23,453 bimanual episodes across 46 tasks.
CAPS keeps the base VLA policy unchanged and spends extra inference only when uncertainty rises. It samples future action chunks with a power distribution and uses Metropolis-Hastings search when entropy crosses a threshold. On RoboTwin 1.0 with π0, average success is 47.4%, compared with 32.2% for π0 and 41.3% for π0 plus TACO. On Simpler-WindowX, it reports 60.5% average success, ahead of π0 and several VLA baselines.