Visual realism checks in simulation-based VLA evaluation
Robotics evaluation teams should add a small visual-cue regression set before using simulation scores to rank VLA policies. VISER gives a practical template: PBR materials, specular highlights, soft shadows, and reconstructed real-world manipulation tasks. The useful check is simple: run the same policy on paired scenes with and without the visual cue, then compare the direction of the result against a real-robot run on a few tasks.
The reason to prioritize this is operational. VISER reports an average Pearson correlation of 0.92 between simulation and real-world policy performance, and its examples show large task-level swings from visual details. In the eggplant-in-pot step, success rises from 10% without specular highlights to 90% with them, close to 100% in the real world. For put-spoon-on-towel, soft shadows give 49% success, compared with 12% without shadows and 42% in the real world. A simulator that drops these cues can rank policies using scenes that remove information the robot uses for geometry and spatial grounding.