UMI-style real-robot evaluation protocol with reset JSON and progress scoring
Teams comparing wrist-view manipulation policies should add a shared physical evaluation pack before treating model scores as comparable. UMI-Bench 1.0 gives a concrete template: fix the workstation, wrist RGB observation setup, action interface, scene reset, rollout logging, and human scoring. Each episode carries a reset image plus scene JSON with task ID, object metadata, position, pose, target region, split, and task-specific factors.
This is useful for labs that see policy rankings move when camera placement, reset procedure, or object pose changes. UMI-Bench reports Full Success Rate and a 0–100 Progress Score across seen and unseen condition cells. Its results show why the split matters: average Progress Score falls from 59.62 in Seen/Seen episodes to 40.19 under combined shifts, and position, layout, pose, or dynamics shifts hurt more than object or appearance changes. Two long-horizon tasks reach 0% Full Success Rate for all three evaluated models, which gives evaluation teams a clear stress test for policies that look acceptable on shorter tasks.