Software-agent research is getting more executable, replayable, and production-bound
Today’s research is strongest where software work can be checked by execution. The main emphasis is stricter evaluation for coding agents, plus better test generation for code and APIs.