Adaptive harnesses and long-horizon evaluation
Test-Time Harness Evolution (TTHE) rewrites and selects executable control programs using unlabeled execution traces. With DeepSeek-V4-Flash, it raised SWE-bench Verified performance from 20.0% to 35.0% and BIRD from 12.0% to 50.0%, while keeping model weights fixed. The method still depends on imperfect proxy signals and can select a weaker harness.
Long-Horizon-Terminal-Bench shows the remaining operational limit. Across 46 containerized tasks, agents averaged 9.9 million tokens and 85.3 minutes per task. The best tested model solved 15.2% at the 0.95 reward threshold, and timeouts caused 79% of unresolved runs. Dense subtask grading made partial progress visible and gave harness designers better failure evidence than a final pass rate alone.