Reliable grounding data
Several papers treat grounding errors as a data problem. SPARC auto-labels robot demonstrations with interacted-object boxes, trajectories, and phase labels, then filters them with a reliability score based on motion, gripper proximity, and robot-body overlap. On IA-Bench, it reports 80.2% interacted-object localization accuracy, compared with 58.1% for a detector-confidence baseline, and keeps 77.6% coverage at a 90% precision operating point.
LabVLA builds laboratory supervision in simulation because standard vision-language-action (VLA) policies rarely see lab instruments, liquids, and protocol workflows. Its RoboGenesis engine creates 2,947 annotated assets, more than 1,000 textures, 10,000 lab scenes, and demonstrations across 16 robot platforms. GIVE adds human gestures to VLA inputs for handover: skeleton overlays, fingertip rays, and short semantic gesture descriptions raise real-world handover success to 80.0%, compared with 0.0% for the base policy in the reported trials.