Trend · Day · 2026-05-14 · Software Intelligence
The day’s strongest signal is practical code-agent work under executable checks. FrontierSmith and DIO-Agent use scoring or execution errors to make coding tasks harder and more useful.
Idea · Day · 2026-05-14 · Software Intelligence
Code-agent work is converging on operational controls that teams can build and test now: sandbox services for rollouts and qualification, per-task permission gates for shell access and skills, and repository context…
Trend · Day · 2026-05-13 · Software Intelligence
The period’s clearest signal: code agents are being judged by complete, checkable work. SWE-Cycle and Phoenix-bench make setup, tests, and domain toolchains part of the score.
Idea · Day · 2026-05-13 · Software Intelligence
Complete agent work now needs evidence that the agent set up the project, chose the right files, ran meaningful checks, and preserved existing behavior.
Trend · Day · 2026-05-11 · Software Intelligence
The strongest signal is that agents need inspectable execution and stricter task evidence. DuST uses execution-labeled candidate code as training data, Shepherd records live agent state for branching, and ComplexMCP…
Idea · Day · 2026-05-11 · Software Intelligence
Agent teams can move faster by adding trace, state, and provenance checks around the places where agents already fail: coding runs, MCP tool workflows, and automation workflows that mix untrusted text with secrets or…
Trend · Week · 2026-W19 · Software Intelligence
This week’s research treats large language model (LLM) coding agents as systems that need proof before trust. The strongest work checks generated code through execution, repository tasks, formal proofs, tool contracts…
Idea · Week · 2026-W19 · Software Intelligence
Coding-agent adoption needs evidence checks inside normal engineering work: a pre-execution verifier for tool calls, repository acceptance rules for maintenance tasks, and staged-ticket security tests for backlog work.
Trend · Day · 2026-05-09 · Software Intelligence
The day’s strongest signal is executable evidence for agent software. Papers test code with generated inputs, diagnose failed runs from telemetry, and enforce contracts around skills or tool actions.
Idea · Day · 2026-05-09 · Software Intelligence
Coding-agent reliability work is converging on small, buildable checks around generated code, failed runs, and reusable skills.
Trend · Day · 2026-05-08 · Software Intelligence
Coding-agent research is testing decision quality under executable checks. FixedBench measures when agents should leave code untouched. SWE Atlas scores everyday repository work.
Idea · Day · 2026-05-08 · Software Intelligence
Teams adopting coding agents can add three concrete controls now: a pre-edit abstention check for stale issues, a repository configuration audit for agent instructions and permissions, and a proof-focused lane for code…