Safety-specific release tests for VLA manipulation policies
VLA manipulation teams should add a safety test stage that measures unsafe contact and unsafe instruction following separately from task completion. LIBERO-Safety gives a usable template: 75 tasks across affordance-aware grasping, human-robot interaction, tabletop spatial avoidance, free-space hand-object avoidance, and semantic safety reasoning, with difficulty levels L0-L2. The benchmark also includes 19,664 human-screened collision-free demonstrations generated from sparse keyposes and CuRobo collision checks.
The operational reason is simple: a policy can finish easy manipulation tasks and still fail around clutter, human hands, or held objects. In LIBERO-Safety, OpenVLA-OFT drops to 1.3% success on hard affordance-aware grasping, and π0.5 reaches only 35.3% on the same level. A practical release gate would run the same policy on standard task suites and on safety suites, log success, collision, refusal, and recovery behavior, and block deployment on tasks where L2 safety cases fail repeatedly.