Research · 2 min read

WEFT Paper Scales Tool Use Post Training Across Whole Agent Systems

An arXiv paper submitted 29 September 2026 argues agent reliability requires evolving the full tool stack, not just model weights, with measured gains on BFCL V4 and long horizon benches.

By Classy AI News · September 30, 2026

WEFT Paper Scales Tool Use Post Training Across Whole Agent Systems

What changed

Researchers posted WEFT (Whole system Evolution For Tool use Post training) to arXiv on 29 September 2026 as paper 2609.36887. The method treats an agent as a coupled system of models, tools, memory, and orchestration, then runs execution driven self evolution across that stack before post training.

Reported results include WEFT 14B beating the Agent World 14B baseline by 6.41 percentage points on BFCL V4, 2.23 points on τ² Bench, and 12.27 points on Claw Eval. A larger WEFT 35B A3B variant is reported on long horizon workflow suites including Toolathlon Verified and AutomationBench.

Research team reviewing benchmark charts on monitors in a lab

Why it matters

Production agents fail when tools drift, state corrupts, or credit assignment spans many turns. Papers that only fine tune the language model leave the harness frozen. WEFT’s claim is that whole system evolution is the scalable path for tool use, with infrastructure like MegaMCP isolating concurrent rollouts over shared services.

For teams shipping agents into finance, ops, or internal IT, the practical question is whether evolution at harness level reduces incident rate faster than prompt patching. The reported bench lifts suggest the approach is competitive, but independent replication on private workflows is still required.

Who is affected

Applied ML leads building tool calling agents should compare WEFT style co evolution against their current RLHF or rejection sampling loops.

MLOps owners running shared MCP or tool gateways should read the MegaMCP isolation design for concurrent rollout safety.

Eval owners should note the paper’s emphasis on prefix preserving sampling and atomic turn credit assignment, which target partial progress retention during long tasks.

What to do next

Reproduce one public benchmark from the paper on your smallest open weight model before committing engineering to a full harness evolution pipeline.

Server racks in a data center supporting large scale model training

What to watch

Whether authors release MegaMCP and evolution logs for third party audit, and any follow up applying WEFT to enterprise MCP deployments with real customer tools.

Sources

  1. Primary. arXiv, WEFT: Scaling Tool Use Post Training for General Purpose Agents (2609.36887) (29 September 2026). Defines method, benchmarks, and reported score deltas.

Newsletter

Get the dispatch

One field. One email when we publish. Privacy.