Clarification checkpoints before irreversible implementation decisions
Coding-agent benchmark designers should evaluate whether an agent asks for a hidden constraint before the implementation decision that depends on it, rather than rewarding question volume or requirement coverage alone. ICAE-Bench finds that recovering more constraints does not automatically raise pass rates, while pAI-Econ-claude concentrates human authority at choices that are costly to reverse. Combining these observations suggests annotating each hidden requirement with its last responsible checkpoint—such as selecting an API, schema, or architecture—and scoring timely clarification separately from final test success and unnecessary interruptions. A small ICAE-Bench replay could compare unrestricted questioning with checkpoint-based access to the simulated user, measuring enhanced-test results and implementation rework.