Analysis · 2 min read

Growing Harness Paper Cuts Agent Inference Cost Up to 98.6 Percent

Shenzhen researchers posted Growing Harness on arXiv in September 2026, showing a learned agent harness can reduce LLM calls by up to 91.8% while keeping strong benchmark success across model sizes from 4B to 120B parameters.

By Classy AI News · September 25, 2026

Growing Harness Paper Cuts Agent Inference Cost Up to 98.6 Percent

What changed

Researchers from the Shenzhen Institutes of Advanced Technology and collaborators submitted Grow the Harness, Not the Context: From Strategy Free Scaffolds to Reusable Specialist Agents to arXiv (2609.26760) in September 2026. The paper introduces Growing Harness, a failure guided training method that turns recurring agent control decisions into persistent executable code while reserving LLM calls for semantic reasoning.

Across BrowseComp Plus and WebArena Verified with deployment models from 4B to 120B parameters, Growing Harness achieved the highest mean success in five of six benchmark model settings. Relative to a tool calling baseline, it reduced LLM calls by 76.0% to 91.8% and deployed agent inference cost by 74.4% to 98.6%. On WebArena Verified, success stayed near 45% across model scales while tool calling fell to 6.7% with the 4B model.

Why it matters

Enterprise agent budgets are dominated by inference tokens, not model list prices. If control logic can migrate from repeated LLM calls into compiled harness code, teams can deploy smaller models without collapsing reliability. That shifts procurement from "buy the biggest frontier model" toward "invest in harness engineering and eval gates."

Who is affected

Platform engineers building internal agent stacks, CFOs approving AI run rate budgets, and vendors selling agent frameworks that default to ReAct style tool loops. Model providers pitching scale alone may face harder ROI questions when harness optimization delivers double digit cost cuts.

What to do next

Instrument one production agent workflow to count LLM calls per task family. If control patterns repeat, pilot a harness growth experiment with held out regression gates before the next model tier upgrade.

What to watch

Open source releases or commercial products claiming Growing Harness style program growth, and whether independent teams replicate the 98.6% cost reduction without success rate collapse on private enterprise tasks.

Software engineers reviewing code architecture on multiple screens
Figure: The method starts from a strategy free scaffold with fixed model and tool interfaces.

The paper's held out gate rolls back harness edits that harm prior capability, addressing a common failure mode in continual agent optimization. That design choice matters for production teams who cannot afford regressions when a harness patch fixes one workflow and breaks three others.

Data center corridor with illuminated server infrastructure
Figure: Cost savings scale with task repetition, not one off demos.

Sources

  1. Primary. arXiv, Grow the Harness, Not the Context: From Strategy Free Scaffolds to Reusable Specialist Agents (September 2026). Establishes Growing Harness method, benchmark results, and cost reduction ranges.

Newsletter

Get the dispatch

One field. One email when we publish. Privacy.