Interactive coding-agent evaluation
Benchmarks are putting the user back into the loop. SWE-Together reconstructed 109 repository-level tasks from 11,260 recorded sessions and scored both final correctness and User Correction, a measure of how much explicit or soft feedback the agent needed. Claude Opus 4.8 led the reported agents at 63% pass@1, while the reference patch baseline stayed around 78%.
SWE-INTERACT made the cost of interaction more visible. On the same underlying tasks, Opus 4.8 dropped from 50.7% single-turn resolve rate to 26.7% in the multi-turn setting. GPT 5.5 dropped from 48.0% to 24.7%, while its per-trial cost rose from $2.78 to $9.84. The failures were not only missed goals; forgotten requirements and implementation bugs remained common after long exchanges.