Explicit grounding gets treated as a control primitive
Grounding is the clearest technical theme. ProGAL-VLA inserts an explicit verification step between language and control: it turns an instruction into a symbolic sub-goal, matches that sub-goal to 3D scene entities, and only then passes a verified goal embedding to the action policy. The payoff is concrete. On LIBERO-Plus robustness it reports 85.5 total, ahead of OpenVLA-OFT+ at 79.6, and robot-perturbation performance rises from 30.3 to 71.5. The same paper also treats ambiguity as a first-class signal, with AUROC 0.81 for ambiguity detection and clarification behavior rising from 0.09 to 0.81. RoboLab reinforces why this matters. Its held-out simulation benchmark shows current generalist policies still fail on most unseen tasks, with π0.5 at 23.3% overall success and only 21.5% on semantic visual grounding.