EVOHARNESSBENCH Shows Tool Growth Can Erase Agent Competence
EVOHARNESSBENCH tests agents when tools, skills, and subagents change over time, finding harness expansion can erase prior competence even as new capabilities arrive.
What changed
Researchers posted EVOHARNESSBENCH to arXiv on 3 September 2026 (2609.04280). The benchmark evaluates LLM agents when the external harness itself evolves across tools, reusable skills, and specialist subagents. It includes 17 multi stage harness streams built from verifier based tasks, covering 802 tasks, 520 tools, 42 skills, and 62 agents.
The authors test two settings: deployment evaluation, which measures whether agents retain competence on earlier tasks as the harness expands, and self evolving adaptation, which asks whether accumulated experience helps when new capabilities arrive. Results show harness expansion alone can degrade performance on previously solved tasks, producing what the paper calls harness induced forgetting.
Why it matters
Most agent evals treat the tool stack as fixed. Production stacks are not fixed. Teams add MCP servers, skills libraries, and router agents weekly. EVOHARNESSBENCH makes non stationarity in the harness a first class metric, which is closer to how platform teams actually ship.
If your roadmap assumes more tools always help, this benchmark is a direct counterexample. Retention and adaptation can pull in opposite directions: preserving old competence does not guarantee faster uptake of new tools.
Who is affected
Agent platform owners adding tools without regression gates. Eval leads who report single pass rates on static SWE benches. Security and compliance teams routing agents through expanding connector catalogs. Model vendors marketing "agentic" upgrades without measuring harness drift.
What to do next
Add a harness evolution slice to your internal eval suite: rerun a frozen task set after each tool or skill addition and track pass rate delta, not just new capability demos.
What to watch
Whether labs adopt EVOHARNESSBENCH alongside static benchmarks like Terminal Bench. Follow up work on mitigation strategies for harness induced forgetting. Vendor claims about self improving agents tested under controlled harness streams rather than marketing demos.
Sources
- Primary. arXiv, EVOHARNESSBENCH: Can Your Agents Keep Pace with an Evolving Harness? (3 September 2026). Benchmark design, scale, and main findings on forgetting versus adaptation.
- Secondary. None required; the arXiv preprint is the sole primary source for numerical claims in this brief.