Harness adaptation gate for small-model workflow migration
Teams running repetitive workflows such as budget approval should test a small language model with a task-specific harness before committing to frontier-model inference costs. The harness optimizer can inspect failed trajectories and edit instructions, tool availability, context selection, hooks, and orchestration loops. In the reported budget-approval task, an adapted Gemma-4-26B-A4B setup reached 98.3% accuracy, up from 75.0% with the default harness and above the 97.3% reported for Gemini-3.1-Pro. Across 21 task-model pairs, adaptation improved 16 and closed the measured gap on seven.
A practical adoption gate is a replay set drawn from one high-volume workflow, split by instance diversity and operational edge cases. Compare the frontier agent, the small model with its current harness, and the adapted small-model agent on task accuracy, contract violations, latency, and cost per completed case. Keep a frontier-model fallback for cases outside the validated workflow. The study found larger gains on repetitive tasks and stronger small models, so a diverse holdout is needed before using the savings estimate for capacity planning.