Physical reasoning must survive execution
IMBench makes the reasoning–action gap measurable. Leading vision-language models reached roughly 74% constraint understanding, but GPT-5.5 completed only 11.3% of tasks from vision and 18.8% with privileged object state. Several alignment, tool-use, hidden-state, and balancing tasks remained at zero.
AC-VLA addresses a related failure at the policy level. It decomposes instructions into reusable sub-tasks and masks wrist views during selected phases to reduce trajectory memorization and visual shortcuts. On LIBERO-OOD, the π₀.₅ variant reached 64.2% Spatial-OOD and 73.3% Goal-OOD success, gains of 28.7 and 26.7 percentage points. Together, the studies support evaluation and training around executable recombination rather than verbal or in-distribution competence alone.