Closed-loop code-agent evaluation
Asuka-Bench evaluates web-app agents under vague initial requests and multi-round user feedback. The benchmark hides the full product requirements, tests rendered browser behavior, and sends direct failure feedback into later rounds. The reported spread is large: after three rounds, weighted task pass rate ranges from 51.8% to 90.1% across 13 model-runtime configurations.
ADK Arena treats Agent Development Kits (ADKs) as measurable engineering choices. One coding agent builds benchmark agents for 51 Python kits inside isolated Docker environments. Generation succeeds in 57% of runs, while cost varies 5.6x across kits. The best single-benchmark agents reach 80% task resolution, and the median kit reaches 32%.