Trend · Week · 2026-W20 · Software Intelligence
Code-agent research this week set a higher bar for useful work. SWE-Cycle and SaaSBench score setup, integration, tests, and delivered behavior. Rollout Cards adds reporting discipline for agent runs.
Idea · Week · 2026-W20 · Software Intelligence
Code-agent work now has enough evidence to support narrower adoption gates: require runnable setup and test proof before accepting agent pull requests, audit evaluation harnesses for score exploits before trusting…
Trend · Day · 2026-05-16 · Software Intelligence
The strongest signal is operational evaluation. 1GC-7RC, AgentKernelArena, and TOBench all score agents inside bounded work loops with tools, runtime checks, and resource limits.
Idea · Day · 2026-05-16 · Software Intelligence
Code agents are ready for narrower operational tests inside engineering workflows: fixed-budget acceptance runs, package-name checks before installs, and scoped code-editing pilots tied to token spend.
Trend · Day · 2026-04-17 · Software Intelligence
The clearest signal for this day is that coding research is tightening the checks that happen before generation, retrieval, or autonomous action.
Idea · Day · 2026-04-17 · Software Intelligence
The clearest short-term builds are control layers that check understanding before an agent acts. The most concrete cases here are a structural repository locator for intent-only queries, a requirement-alignment gate…
Trend · Day · 2026-04-07 · Software Intelligence
The strongest work on this day makes software agents easier to constrain, inspect, and score. CodeStruct and SWE-Shield tighten code-agent evaluation around exact edits and design rules.
Idea · Day · 2026-04-07 · Software Intelligence
Concrete changes are showing up in three places: repository repair agents can work on named code entities and be judged on small valid diffs, patch evaluation needs design-constraint checks beyond test pass rate, and…