Broader coding-agent evaluation
ICAE-Bench tests whether an agent can clarify incomplete requirements and build a repository, rather than solve a fully specified edit. Its 480 tasks span 12 languages; six models across two harnesses still struggled with hidden constraints, boundary cases, and long-horizon integration.
Tencent WorkBuddy Bench extends evaluation across code, web, office, and security work. Its 260 tasks are reverse-engineered and rewritten to reduce recovery through web search, then released with environments and verifiers for auditability. The suite deliberately avoids a single aggregate score because each domain uses a different verification method. Together, the benchmarks make task construction and evaluator design part of the capability claim, not background implementation detail.