Agent work is being judged by complete loops, visible failures, and auditable evidence
The day’s strongest signal is practical measurement of agent work under real constraints. MAC and TeleSWEBench show limited autonomy in agent design and domain code repair.