Code agents are being graded on complete, verifiable work
The period’s clearest signal: code agents are being judged by complete, checkable work. SWE-Cycle and Phoenix-bench make setup, tests, and domain toolchains part of the score.
The period’s clearest signal: code agents are being judged by complete, checkable work. SWE-Cycle and Phoenix-bench make setup, tests, and domain toolchains part of the score.
Complete agent work now needs evidence that the agent set up the project, chose the right files, ran meaningful checks, and preserved existing behavior.