Harness-dependent evaluation
AgentCompass separates benchmarks, harnesses, and execution environments, then shows that the same model’s result can vary with the harness. On SWE-bench-Pro, Claude-Opus-4.8 scored 66.21 with Mini-SWE-agent and 73.87 with OpenHands. A separate set-shifting benchmark finds that agents quickly settle into recurring tool routines and may keep using an unreliable group after a silent backend change. Together, the studies make interaction history and harness configuration part of the measured capability, rather than incidental test plumbing.