Abstention and engineering-quality benchmarks
FixedBench targets a failure mode that normal issue-resolution scores miss: agents edit code after the correct patch has already been applied. In the main resolved setting, agents still made unwanted executable-code edits in 35% to 65% of cases. A direct “Abstain or Fix” prompt improved abstention for some models, but it also caused harmful inaction on partially fixed issues.
SWE Atlas widens evaluation to codebase Q&A, test writing, and refactoring across 18 active repositories. Its rubric checks expose quality gaps after functional tests pass. Top native-scaffold results stay near the low 40s in Pass@1, and best Pass³ values reach only 29.2, so consistency remains a clear bottleneck.