Benchmarks are being pressed to measure precise work, not just passing outputs
Evaluation is getting stricter about what counts as a real fix. PDB shows that high unit-test scores can hide poor debugging behavior when models rewrite too much code. On PDB-Single-Hard, GPT-5.1-Codex reaches 76.1% unit-test score but only 39.7% edit precision, while Claude-4.5-Sonnet reaches 71.8% precision and Gemini-2.5-Pro 71.4%. The same pattern holds on multi-bug programs, where top unit-test scores still come with weak precision. A companion signal comes from app-building: in a single React Native task on one GH200, SWE-Bench standings did not predict the best shipped result. Kimi-K2.5 Q3 was the only model judged fully spec-compliant at the app level, while GLM-5.1 and DeepSeek-V3.2 failed to run out of the box for concrete operational reasons.