Local real-robot regression bench for VLA policy releases
A VLA lab can make a small physical benchmark part of every policy release: rebuild the same SO-101 arm setup, run a fixed set of manipulation tasks, and publish per-task success on in-distribution and out-of-distribution scenes. VLA-REPLICA gives enough detail for this workflow to be practical: an SO-101 6-DoF arm, RealSense D455 top camera, wrist webcam, 32 inch light box, commodity objects, camera overlay tools, AprilTag alignment, fixed lighting, reference images, and predefined placements. The reported setup cost is about $1050, and a new user assembled it in under an hour.
The useful check is simple: train or fine-tune on the same 500 expert demonstrations, run the 90 defined test scenes, and compare against reported baselines such as π₀.₅ at 0.54 average success on the 10 in-distribution tasks. This gives robotics groups a cheap physical gate before claiming that a VLA policy works outside simulation.