Trend

Agent evaluation reaches ambiguous projects as reliability moves into the harness

Day · 2026-07-23 · Software Intelligence

After several days centered on executable feedback inside coding loops, today’s evidence broadens the control surface. New benchmarks test agents on incomplete product intent and mixed workplace tasks, while reliability mechanisms deliver memory, logic, and review at defined checkpoints. Results remain early: several studies lack broad quantitative comparisons, and one workflow improves auditability at substantial cost.

Broader coding-agent evaluation

ICAE-Bench tests whether an agent can clarify incomplete requirements and build a repository, rather than solve a fully specified edit. Its 480 tasks span 12 languages; six models across two harnesses still struggled with hidden constraints, boundary cases, and long-horizon integration.

Tencent WorkBuddy Bench extends evaluation across code, web, office, and security work. Its 260 tasks are reverse-engineered and rewritten to reduce recovery through web search, then released with environments and verifiers for auditability. The suite deliberately avoids a single aggregate score because each domain uses a different verification method. Together, the benchmarks make task construction and evaluator design part of the capability claim, not background implementation detail.

Reliability as enforced infrastructure

Three systems place critical controls outside the model’s voluntary behavior. Cue-anchored working memory injects scoped facts when files, symbols, or lifecycle events trigger them; the agent made no memory calls in 114 turns under the strongest voluntary control, while harness delivery survived repeated compaction. Euclid-MCP delegates rule deduction to Prolog and returns proof traces, although its reported performance advantage lacks numerical baseline results.

For economic theory, pAI-Econ-claude uses inspectable intermediate records, targeted gates, and human checkpoints where no automatic oracle exists. Blinded evaluators preferred it in four of five matched tasks, but it consumed 4.6 to 18 times the baseline usage allowance. The common finding is bounded: enforced delivery and explicit gates improve traceability and error interception, but they do not establish cheap or general correctness.

OlderExecutable interfaces are becoming the common lever for robot reliability