Auditable agent demos
Google’s claim that agents built an operating system for about $900 became a case study in missing evidence. The critique does not rerun the task. It asks for the artifacts needed to judge the result: the full prompt, generated source code, run logs, retry counts, dry-run history, and checks for copied public code.
The reported setup also matters. The task used specialized roles, subagent delegation, restart infrastructure, and an anti-cheating component. Those pieces may be valid engineering, but they make the result hard to attribute to model capability alone without logs and release artifacts.