Benchmark integrity
Agent benchmark numbers receive direct adversarial pressure in this period. BenchJack audited 10 agent benchmarks and generated working reward-hacking exploits for all 10, with 219 distinct flaws across its taxonomy. Its patching loop reduced hackable-task ratios below 10% on four fixable benchmarks, and fully patched WebArena and OSWorld within three iterations.
Rollout Cards addresses a related reporting problem. The paper argues that agent studies need rollout records, declared scoring views, reporting rules, and omitted-field manifests. In its audit of 50 popular repositories, none reported failed, errored, or skipped rollouts alongside headline scores. Re-grading fixed artifacts changed reported scores by as much as 20.9 points and could swap model rankings on tau-bench.