Research · 2 min read

Stateless Language Agents Keep Research State Outside the Chat Window

A new arXiv paper argues long horizon research agents fail when conversation history becomes the system of record, and proposes stateless workers with a harness owned search state.

By Classy AI News · October 8, 2026

Stateless Language Agents Keep Research State Outside the Chat Window

What changed

Researchers posted Stateless Language Agents: Scaling Long Horizon Automated Research on arXiv on 7 October 2026 (identifier 2610.07625). The paper targets automated research systems that run language model agents for long horizons. It argues that adding inference alone does not guarantee progress: agents replay growing histories, duplicate work, or stop experimenting while token use continues.

The authors attribute those failures to where research state lives and who assigns the next experiment. They introduce Stateless Language Agents (SLAs) built on stateful search with stateless agents: no agent carries its conversation across invocations. Instead a harness owns candidate solutions and measured outcomes, then rebuilds a fresh, role specific context on every call.

Their SLA framework uses a stateless Advisor that reads harness summarized evidence across search directions and assigns concrete experiments to parallel Workers. They report evaluations against three recent frameworks on software engineering, kernel optimization, and algorithm design at budgets up to one billion tokens, claiming SLA reaches the best final result on every task and matches a strong kernel baseline final performance with more than 84 percent fewer tokens.

Researchers working at desks with multiple monitors

Why it matters

Teams building autonomous R&D copilots often treat chat logs as memory. This paper is a design critique with benchmarks at extreme token budgets, not a product launch. If the harness first pattern holds, platform owners should separate durable experiment state from ephemeral model calls before they scale overnight research jobs.

The Advisor consuming less than 0.6 percent of tokens in their ablations also challenges teams that spend most compute re summarizing old messages.

Who is affected

Applied AI leads running agentic research or code optimization, eval owners designing long horizon benchmarks, and vendors marketing autonomous scientists without persistent state tooling.

What to do next

Before expanding autonomous research pilots, map where hypotheses, metrics, and failed runs are stored today; if the answer is in the thread, pilot a harness owned state store.

What to watch

Independent reproduction on the paper’s kernel and algorithm design tasks, plus open source releases if authors ship the SLA framework beyond the arXiv description.

Data visualization charts on a laptop screen

Sources

  1. Primary. arXiv, Stateless Language Agents: Scaling Long Horizon Automated Research (7 October 2026). Defines SLA architecture, benchmarks, and reported token savings.
  2. Secondary. Search arXiv mirror, 2610.07625 listing (7 October 2026). Confirms submission date and abstract text.

Newsletter

Get the dispatch

One field. One email when we publish. Privacy.