Execution discipline in coding agents
Coding-agent papers keep adding execution discipline around long tasks. KISS Sorcar uses forced continuation summaries, tool access, and git worktree isolation to keep repo edits reviewable and recoverable. It reports a 62.2% pass rate on Terminal Bench 2.0 with Claude Opus 4.6, slightly above Claude Code at 58% and Cursor Composer 2 at 61.7%. The details matter: only 43.8% of tasks pass in all five trials, and 19 tasks always fail. The result is better repo handling, not stable autonomy.