Contact-region training and scoring for manipulation VLA policies
Manipulation teams should add contact-region checks to VLA evaluation when the task depends on object parts, such as handles, lids, buttons, and tool tips. The operational failure is simple: the model can identify the right object and still touch the wrong part.
AffordVLA gives a practical training route. It uses a frozen affordance teacher during training to align intermediate VLA visual tokens with task-conditioned affordance features, then removes the teacher at inference. That keeps the deployed policy path unchanged while pushing the visual representation toward functional interaction regions. A cheap first test is a held-out set of part-sensitive tasks with masks or sparse human labels for the intended contact area, scored by both task success and first-contact accuracy. The contact score should catch failures that a coarse success metric can hide.