Research · 2 min read

Agent Harness Plans Beat Sham Guidance by Seven Points on Retail Tasks

An arXiv study of stateful LLM agents on tau bench finds task specific plans raise oracle verified success by 7.17 percentage points, while a read only terminal verifier blocks most invalid completions at under one cent per episode.

By Classy AI News · September 18, 2026

Agent Harness Plans Beat Sham Guidance by Seven Points on Retail Tasks

What changed

Researchers Yukun Zhang, Kemu Xu, and Yishen Chen posted How Do Agent Harnesses Create Value? Planning Information and Release Control in Stateful LLM Agents to arXiv on 17 September 2026 (arXiv:2609.20474). The team compared Fixed task specific plans against Sham guidance with matched word count across 265 matched cells on Retail and Airline style tasks in tau bench.

Fixed planning raised oracle verified success by 7.17 percentage points (90% bootstrap interval 1.15 to 13.36 points), with the largest gains on higher complexity tasks. A read only terminal verifier rejected 61% of oracle invalid Retail episodes while withholding 17% of correct ones, adding less than one cent of cost per episode.

Software developer reviewing code on screen

Why it matters

Enterprise agent rollouts often budget model upgrades while treating harness design as plumbing. This paper quantifies where value actually sits: planning information when false acceptance is cheap, and verification when false passes are expensive. The authors report a standalone verifier captures nearly all false pass reduction of the full planning plus verification stack at a fraction of cost.

For procurement, the implication is to score vendors on harness artifacts (plans, verifiers, release gates) rather than parameter counts alone. Teams running customer service or operations agents on stateful tool APIs should expect Sham level guidance (policy text without task structure) to underperform even when word counts match.

Who is affected

Applied AI platform engineers choosing between Codex style harnesses, custom orchestrators, and packaged agent frameworks. Risk and compliance owners who need deterministic rejection of invalid tool outcomes without blocking every edge case. Benchmark consumers comparing tau bench scores across labs where harness details differ. Investors funding agent startups should ask whether moats live in models or in planning and verification stacks.

What to do next

Audit one production agent workflow: document whether plans are task specific or generic policy dumps, and measure false completion rate with a read only terminal check. If liability from wrong completions is high, prioritize verifier deployment before larger model spend.

What to watch

Follow whether tau bench maintainers publish harness disclosure requirements. Watch for replication on airline and telecom domains beyond Retail pilots. Track if OpenAI, Anthropic, or Microsoft productize standalone verifiers aligned with the paper's cost claims.

Whiteboard with workflow diagrams in a tech office

Sources

Newsletter

Get the dispatch

One field. One email when we publish. Privacy.