Mechanism-aware long-horizon evaluation
LongBench gives this period a concrete anchor: long-horizon robot evaluation is getting more specific about why policies fail. The benchmark covers 10 real-world tasks and more than 1,000 episodes. It separates fully observable execution problems from context-dependent ambiguity, then scores progress stage by stage instead of using a single pass/fail number. That matters because current policies break for different reasons. On context-independent tasks, pi_0 leads with an average stage-wise score of 86.3, while Diffusion Policy reaches 51.2 and OpenVLA-OFT 32.7. Dynamic grasping is the clearest stress case. pi_0 gets 73.3, while the other listed systems stay between 0.0 and 13.3.