Research · 2 min read

NeoHorse 1 Closes an Evaluation Selection Update Loop for Agent Native Models

A 8 September arXiv paper introduces NeoHorse 1, an agent native model family that records routing decisions and converts them into validated training data, raising macro average scores on eleven agent benchmarks.

By Classy AI News · September 9, 2026

NeoHorse 1 Closes an Evaluation Selection Update Loop for Agent Native Models

What changed

Researchers posted NeoHorse 1 on arXiv on 8 September 2026 (identifier 2609.08183). The system pairs a heterogeneous model pool with a routing harness that logs predicted capability demand, selected service tier, and subsequent tool interactions for each user turn. Those logs become training examples that preserve interleaved reasoning, tool calls, and harness context, then pass structural validation, six dimensional semantic evaluation, and subscene labeling before entering fine tuning.

The authors report macro average gains on eleven benchmarks spanning harness based agents, tool use, coding, and instruction following: from 58.94 to 64.87 at 4B parameters and from 65.60 to 69.04 at 9B parameters, narrowing the gap between the post trained 4B model and the 9B base model.

Why it matters

Most agent stacks treat routing and evaluation as plumbing. NeoHorse 1 treats routing telemetry as the curriculum. For applied AI teams building multi model routers, that means production traffic can feed the next training mixture instead of relying only on static offline sets. The paper is an early prototype of harness mediated recursive self improvement, not a claim of fully autonomous model rewriting.

Who is affected

Platform teams operating model gateways, agent harness vendors, and enterprise AI groups running tiered model pools should read the routing guided on policy distillation section. Investors tracking post training infrastructure should note the explicit evaluation selection update loop as a design pattern competitors may copy.

What to do next

If you operate a router, export one week of anonymized routing decisions with outcome labels and ask whether your current fine tuning pipeline can ingest them without manual relabeling. If not, NeoHorse 1 defines a target architecture.

What to watch

Whether NeoHorse 1 authors release open weights or a public harness API, and whether independent groups reproduce the 4B to 9B gap closure on private enterprise benchmarks.

Sources

  1. Primary. Zehua Pei et al., NeoHorse 1: Towards Recursive Self Improvement via Agentic Post Training with Routing Harness (8 September 2026). Full method, benchmark table, and routing loop description.
  2. Secondary. arXiv cs.CL September 2026 listing, 2609.08183 submission record (8 September 2026). Confirms submission date and subject classification.

Newsletter

Get the dispatch

One field. One email when we publish. Privacy.