Release tests for tool failover within persistent agent sessions
Agent-platform teams operating redundant APIs should add silent backend shifts to release testing and run them across every supported harness. AgentCompass shows that changing the harness can move the same model’s SWE-bench-Pro score by several points and can expose different trajectory failures. The set-shifting benchmark shows a separate operational risk: after reliability changes during a persistent session, agents can remain locked into an obsolete tool routine even when an equivalent working tool is available. A release report should therefore include a model-by-harness matrix with post-shift tool shares, task completion, repeated calls, and recovery latency—not only aggregate task success. The cheapest check is to replay a small set of production-like sessions, silently fail the preferred backend halfway through, and compare recovery with a fresh-session control; this determines whether session history or the backend failure itself causes the loss.