Growing Harness Paper Cuts Agent Inference Cost Up to 98.6 Percent
Shenzhen researchers posted Growing Harness on arXiv in September 2026, showing a learned agent harness can reduce LLM calls by up to 91.8% while keeping strong benchmark success across model sizes from 4B to 120B parameters.
What changed
Researchers from the Shenzhen Institutes of Advanced Technology and collaborators submitted Grow the Harness, Not the Context: From Strategy Free Scaffolds to Reusable Specialist Agents to arXiv (2609.26760) in September 2026. The paper introduces Growing Harness, a failure guided training method that turns recurring agent control decisions into persistent executable code while reserving LLM calls for semantic reasoning.
Across BrowseComp Plus and WebArena Verified with deployment models from 4B to 120B parameters, Growing Harness achieved the highest mean success in five of six benchmark model settings. Relative to a tool calling baseline, it reduced LLM calls by 76.0% to 91.8% and deployed agent inference cost by 74.4% to 98.6%. On WebArena Verified, success stayed near 45% across model scales while tool calling fell to 6.7% with the 4B model.
Why it matters
Enterprise agent budgets are dominated by inference tokens, not model list prices. If control logic can migrate from repeated LLM calls into compiled harness code, teams can deploy smaller models without collapsing reliability. That shifts procurement from "buy the biggest frontier model" toward "invest in harness engineering and eval gates."
Who is affected
Platform engineers building internal agent stacks, CFOs approving AI run rate budgets, and vendors selling agent frameworks that default to ReAct style tool loops. Model providers pitching scale alone may face harder ROI questions when harness optimization delivers double digit cost cuts.
What to do next
Instrument one production agent workflow to count LLM calls per task family. If control patterns repeat, pilot a harness growth experiment with held out regression gates before the next model tier upgrade.
What to watch
Open source releases or commercial products claiming Growing Harness style program growth, and whether independent teams replicate the 98.6% cost reduction without success rate collapse on private enterprise tasks.

The paper's held out gate rolls back harness edits that harm prior capability, addressing a common failure mode in continual agent optimization. That design choice matters for production teams who cannot afford regressions when a harness patch fixes one workflow and breaks three others.

Sources
- Primary. arXiv, Grow the Harness, Not the Context: From Strategy Free Scaffolds to Reusable Specialist Agents (September 2026). Establishes Growing Harness method, benchmark results, and cost reduction ranges.