Research · 2 min read

ActiveSaddler Co Evolves Agent Harness Curricula From Failure Patterns

Microsoft researchers posted ActiveSaddler on arXiv 2610.00906, framing harness optimization as a non stationary bandit over failure patterns. GAIA2 and Terminal Bench 2.0 Pass at 1 gains reached 4.4 and 7.5 points.

By Classy AI News · October 4, 2026

ActiveSaddler Co Evolves Agent Harness Curricula From Failure Patterns

What changed

Researchers published ActiveSaddler on arXiv 2610.00906, treating automated agent harness optimization as an automated curriculum learning problem. Instead of fixing which training scenarios generate feedback while prompts and tool interfaces evolve, ActiveSaddler models the curriculum as a non stationary bandit whose arms are reusable failure patterns extracted from execution traces.

The paper reports gains on GAIA2 and Terminal Bench 2.0: 4.4 and 7.5 percentage point improvements in test Pass at 1 versus the same harness optimizer with a scenario order fixed before optimization. Code and a project site are promised at `https://aka.ms/ActiveSaddler-website`.

Researchers reviewing agent evaluation dashboards on large monitors

Why it matters

Teams shipping coding and tool using agents already iterate prompts, parsers, and control logic from logs. ActiveSaddler argues the missing dimension is which failures to revisit next as the harness improves. That matters for applied AI leads who budget eval cycles: static scenario lists can leave easy wins on the table once early fixes land.

Who is affected

Agent platform owners, eval engineers, and internal tooling teams running harness bakeoffs on benchmarks like GAIA style suites or terminal coding tasks. Investors underwriting agent startups should treat curriculum co evolution as a differentiator in reliability roadmaps, not just raw model upgrades.

What to do next

If you maintain an agent harness, log failures with stable pattern tags (tool timeout, bad JSON, wrong file path) and test whether your next optimization sprint targets the highest estimated learning progress arms rather than a fixed regression suite.

What to watch

Whether ActiveSaddler code drops with reproducible configs, and whether independent groups replicate the 4.4 / 7.5 point lifts on held out enterprise tasks beyond the paper benchmarks.

Software engineers pair programming beside a whiteboard of agent workflows

Sources

  1. Primary. arXiv, ActiveSaddler: Automated Curriculum Learning for Agent Harness Optimization (2610.00906) (October 2026). Defines the bandit formulation, failure pattern arms, and reported Pass at 1 lifts on GAIA2 and Terminal Bench 2.0.
  2. Secondary. ArXiv Signals index entry for 2610.00906 (October 2026). Corroborates title, authorship listing, and benchmark names.

Newsletter

Get the dispatch

One field. One email when we publish. Privacy.