ActiveSaddler Co Evolves Agent Harness Curricula From Failure Patterns
Microsoft researchers posted ActiveSaddler on arXiv 2610.00906, framing harness optimization as a non stationary bandit over failure patterns. GAIA2 and Terminal Bench 2.0 Pass at 1 gains reached 4.4 and 7.5 points.
What changed
Researchers published ActiveSaddler on arXiv 2610.00906, treating automated agent harness optimization as an automated curriculum learning problem. Instead of fixing which training scenarios generate feedback while prompts and tool interfaces evolve, ActiveSaddler models the curriculum as a non stationary bandit whose arms are reusable failure patterns extracted from execution traces.
The paper reports gains on GAIA2 and Terminal Bench 2.0: 4.4 and 7.5 percentage point improvements in test Pass at 1 versus the same harness optimizer with a scenario order fixed before optimization. Code and a project site are promised at `https://aka.ms/ActiveSaddler-website`.

Why it matters
Teams shipping coding and tool using agents already iterate prompts, parsers, and control logic from logs. ActiveSaddler argues the missing dimension is which failures to revisit next as the harness improves. That matters for applied AI leads who budget eval cycles: static scenario lists can leave easy wins on the table once early fixes land.
Who is affected
Agent platform owners, eval engineers, and internal tooling teams running harness bakeoffs on benchmarks like GAIA style suites or terminal coding tasks. Investors underwriting agent startups should treat curriculum co evolution as a differentiator in reliability roadmaps, not just raw model upgrades.
What to do next
If you maintain an agent harness, log failures with stable pattern tags (tool timeout, bad JSON, wrong file path) and test whether your next optimization sprint targets the highest estimated learning progress arms rather than a fixed regression suite.
What to watch
Whether ActiveSaddler code drops with reproducible configs, and whether independent groups replicate the 4.4 / 7.5 point lifts on held out enterprise tasks beyond the paper benchmarks.

Sources
- Primary. arXiv, ActiveSaddler: Automated Curriculum Learning for Agent Harness Optimization (2610.00906) (October 2026). Defines the bandit formulation, failure pattern arms, and reported Pass at 1 lifts on GAIA2 and Terminal Bench 2.0.
- Secondary. ArXiv Signals index entry for 2610.00906 (October 2026). Corroborates title, authorship listing, and benchmark names.