Coding-agent research is being judged by runnable proof, repo realism, and harness quality
This week’s coding-agent research is strongest when claims end in runnable evidence. Benchmarks and systems keep asking whether code builds, executes, and survives workflow checks.