Agent Harness Plans Beat Sham Guidance by Seven Points on Retail Tasks
An arXiv study of stateful LLM agents on tau bench finds task specific plans raise oracle verified success by 7.17 percentage points, while a read only terminal verifier blocks most invalid completions at under one cent per episode.
What changed
Researchers Yukun Zhang, Kemu Xu, and Yishen Chen posted How Do Agent Harnesses Create Value? Planning Information and Release Control in Stateful LLM Agents to arXiv on 17 September 2026 (arXiv:2609.20474). The team compared Fixed task specific plans against Sham guidance with matched word count across 265 matched cells on Retail and Airline style tasks in tau bench.
Fixed planning raised oracle verified success by 7.17 percentage points (90% bootstrap interval 1.15 to 13.36 points), with the largest gains on higher complexity tasks. A read only terminal verifier rejected 61% of oracle invalid Retail episodes while withholding 17% of correct ones, adding less than one cent of cost per episode.
Why it matters
Enterprise agent rollouts often budget model upgrades while treating harness design as plumbing. This paper quantifies where value actually sits: planning information when false acceptance is cheap, and verification when false passes are expensive. The authors report a standalone verifier captures nearly all false pass reduction of the full planning plus verification stack at a fraction of cost.
For procurement, the implication is to score vendors on harness artifacts (plans, verifiers, release gates) rather than parameter counts alone. Teams running customer service or operations agents on stateful tool APIs should expect Sham level guidance (policy text without task structure) to underperform even when word counts match.
Who is affected
Applied AI platform engineers choosing between Codex style harnesses, custom orchestrators, and packaged agent frameworks. Risk and compliance owners who need deterministic rejection of invalid tool outcomes without blocking every edge case. Benchmark consumers comparing tau bench scores across labs where harness details differ. Investors funding agent startups should ask whether moats live in models or in planning and verification stacks.
What to do next
Audit one production agent workflow: document whether plans are task specific or generic policy dumps, and measure false completion rate with a read only terminal check. If liability from wrong completions is high, prioritize verifier deployment before larger model spend.
What to watch
Follow whether tau bench maintainers publish harness disclosure requirements. Watch for replication on airline and telecom domains beyond Retail pilots. Track if OpenAI, Anthropic, or Microsoft productize standalone verifiers aligned with the paper's cost claims.
Sources
- Primary — arXiv, How Do Agent Harnesses Create Value? Planning Information and Release Control in Stateful LLM Agents (17 September 2026). Fixed vs Sham planning results, verifier rejection rates, and cost per episode.
- Secondary — Classy AI News, Tau Tau Bench Shows Coding Agents Fail Real Client Style Agent Builds (15 September 2026). Prior tau bench context on agent construction difficulty.