Full-stack delivery benchmarks
SaaSBench makes enterprise software delivery the test. Its tasks include long product requirements, Docker runtimes, dependency-ordered validation nodes, multiple languages, databases, and frontend/backend stacks. The best reported Pass@1 is 20.68%, and over 95% of failures happen before deep business logic, often during setup, configuration, integration, premature stopping, or stalled debugging.
WebGameBench gives the same point a user-visible form. Agents build browser games, then a runtime evaluator controls Chrome through Playwright and checks behavior against the requirement. The best configuration reaches 76.9% usable rate, but only 20.2% excellent rate. A page can load and still miss rules, input handling, scoring, restart behavior, or win/loss conditions.