Executable repository training and long-run evaluation
Repository-level coding agents need rebuildable projects, real tests, and long trajectories. KAT-Coder-V2.5 reports that AutoBuilder raised executable environment construction success from 16.5% to 57.2% and produced more than 100,000 verifiable environments across 12 languages. Its training pipeline also filters trajectories by exploration, localization, patch quality, verification, recovery, and honesty.
Long-run evaluation follows the same emphasis. EdgeBench measures agents over 134 real-world tasks, including 36 systems and software engineering tasks, with about 38,000 hours of interaction. It reports that 12-hour learning curves fit a log-sigmoid form with mean R² = 0.998. EvoAgentBench adds a procedure-transfer test: every test task shares at least one verified Ability with a training task, yet automatic memory methods still show mixed gains and some large negative-transfer cases.