Evidence packages for long-running coding-agent demos
AI teams running large coding-agent demos should publish a compact evidence package with each claim: prompt versions, generated source code, full agent logs, retry and dry-run counts, restart events, human approvals, total tokens, dollar cost, and a similarity report against public code. Google’s operating-system demo is a clear test case. The public writeup reported $916.92 in API fees and a 2.6B-token budget, while the critique found no released long prompt, source code, or run logs, and no similarity or log analysis for copied public toy operating-system code.
This is a buildable evaluation layer. An internal release gate could require the bundle before a demo is used in a launch post or sales proof. External readers would still need judgment, but they could inspect whether the result came from a general agent run, task-specific scaffolding, repeated restarts, or unreported prompt work.