Repository context and workplace delivery
Agent evaluation is moving toward the conditions that decide whether work ships: the right files, usable artifacts, preserved state, and measurable cost. DeepDiscovery shows that repository tasks need connected context across code, configuration, tests, and organizational structure. Its reported SWE-bench Verified solve rate reaches 78.6%, 8.2 points above its baseline.
EnterpriseClawBench adds the workplace side. It turns real enterprise sessions into 852 reproducible tasks with fixtures, deliverables, hard rules, traces, runtime, token use, and cost. The best audited Lite result is 0.663, which leaves plenty of room on artifact quality and delivery. The open-source census adds scale: agent traces across more than 180 million repositories require multiple detection signals, since pull requests, commits, author patterns, and config files capture different populations.