Evaluation is moving past final-answer scoring
Benchmarks on this day keep asking a stricter question: does a model capture program intent, and can it explain that intent with checkable evidence. CodeSpecBench tests executable preconditions and postconditions, and the best repository-level pass rate is 20.2% on 500 SWE-bench Verified issues. CodeRQ-Bench then grades the reasoning itself. Its VERA evaluator beats prior judges across generation, summarization, and classification tasks, with gains up to 0.26 AUCROC. The message is practical: output quality alone still hides large semantic and reasoning gaps.