Production evaluation for coding agents
The strongest work today tries to grade coding agents on signals that can survive real engineering conditions. ProdCodeBench builds tasks from actual developer-assistant sessions in an industrial monorepo, keeps the original prompts, backs out the landed diff, and scores agents with stable fail-to-pass and pass-to-pass tests. That gives offline evaluation more of the constraints teams face in practice. ClickHouse adds a field report from the deployment side: agents are useful when they can read code, run tools, and stay inside a review-and-test loop. The numbers there are operational, not benchmark-grade, but they show why production teams care about grounded evaluation at all.