Occlusion and physical benchmarks expose brittle manipulation results
Several papers tighten evaluation for vision-language-action (VLA) robot policies, where VLA means mapping language and visual input to robot actions. LIBERO-Occ adds 2,000 occluded LIBERO tasks and reports large drops when task objects or receptacles are hidden. VIM recovers part of the loss by generating a complementary wrist or gripper view, reaching 65.05% average success without a real extra view.
UMI-Bench 1.0 gives Universal Manipulation Interface policies a real-robot tabletop protocol with fixed resets, wrist-view inputs, rollout logging, and human scoring. Its results show that physical factors such as layout, pose, and dynamics hurt more than appearance or object-category changes. A separate sim-real study finds REALM tracks real robot policy rankings better than VLA-Arena and SIMPLER, with Spearman correlation 0.700 before simulator post-training and 0.875 after it.