Trend · Day · 2026-04-25 · Software Intelligence
April 25’s coding research is strongest where claims meet executable evidence. Simulating and Evaluating Agentic Systems and CUJBench both insist on judging agents through full runs, tool traces, state changes, and…
Idea · Day · 2026-04-25 · Software Intelligence
Executable evidence is becoming the practical standard for both agent evaluation and coding workflows. The clearest near-term builds are a replayable evaluator that checks real state and tool traces, a repository intake…
Trend · Day · 2026-04-16 · Software Intelligence
This period centers on coding agents that get better by compressing evidence, pruning weak trajectories early, and testing themselves in harder environments.
Idea · Day · 2026-04-16 · Software Intelligence
Coding-agent work in this window supports three concrete changes: add trajectory compression before reruns on repository tasks, add mid-run budget control for small-model agents, and evaluate production-facing agents in…