End-to-end software delivery benchmarks
Benchmarks treated coding as a full work cycle. SWE-Cycle asks agents to set up repositories, change code, and write verification tests across 489 GitHub issue instances. Phoenix-bench adds hardware-engineering repositories and executable EDA checks, where agents transfer poorly because domain toolchains and project structure matter. SaaSBench and WebGameBench push the same standard into delivered applications: enterprise SaaS systems and playable browser games are judged by runtime behavior, configuration, and cross-component integration.