Trend · Week · 2026-W17 · Software Intelligence
This week’s coding-agent research is strongest when claims end in runnable evidence. Benchmarks and systems keep asking whether code builds, executes, and survives workflow checks.
Idea · Week · 2026-W17 · Software Intelligence
Coding-agent work this week points to three practical changes: treat repository setup as its own executable stage, evaluate repo generation with size-aware runnable tests, and put harness features under explicit…
Trend · Day · 2026-04-22 · Software Intelligence
April 22’s research is strongest where coding work meets concrete checks from real use, project scaffolding, and direct execution. SWE-chat grounds coding-agent claims in kept code and user pushback.
Idea · Day · 2026-04-22 · Software Intelligence
Coding-agent evaluation is moving closer to what teams can verify in their own repositories and pipelines. The most usable directions here are a commit-linked scorecard for kept code and review friction, a narrow…