Executable evidence is the main standard
Evaluation kept tightening around executable proof. Daily trend documents across the week repeatedly favor systems scored by full runs, tool traces, state changes, and live workflow survival. The item-level evidence adds the same message at repository scale: RealBench tests code generation with real repositories, UML design inputs, and human-verified tests, and the best average Pass@1 is still 19.39%. The week reads as a demand for runnable artifacts, not polished code text.