Research · 2 min read

Agensh Shows Agent Count Scales Coding Reliability Without a Central Orchestrator

Microsoft researchers report a decentralized multi agent harness scaling to 1,024 workers, raising ProgramBench pass rates by about 49% at 128 agents on GPT 5.6 Sol.

By Classy AI News · September 23, 2026

Agensh Shows Agent Count Scales Coding Reliability Without a Central Orchestrator

What changed

Researchers from Microsoft posted Agensh: Scaling Organizational Intelligence to 1,024 Agents to arXiv on 23 September 2026 (arXiv:2609.26781). The harness removes a central orchestrator. Concurrent workers run a cooperation loop: gather context, claim subtasks, act, verify, and merge progress asynchronously through a shared workspace, message interface, and shared context store.

Researchers collaborating around whiteboards covered with system diagrams
Figure: Large agent organizations need coordination primitives, not only bigger single models.

On the five hardest ProgramBench tasks evaluated with GPT 5.6 Sol (high), scaling from 1 to 128 agents raised the mean final test pass rate from 19.31% to 28.78%, about a 49% relative improvement, the authors report. On the pandoc task, scaling from 1 to 1,024 agents raised the final test pass rate from 33.89% to 55.06%. Worker trajectories show self organized cooperation patterns that standardize as organizations grow.

Why it matters

Platform teams hitting latency walls on coding agents now have peer reviewed evidence that agent count is its own scaling axis. A harness that avoids a single orchestrator bottleneck matters for anyone building internal agent factories under hard deadlines. The paper also documents emergent coordination without hand authored workflows, which shifts eval focus from solo model scores to organization level reliability.

Engineers monitoring distributed job queues on a wall of screens
Figure: Async worker pools resemble microservice fleets more than chat threads.

Who is affected

Applied AI leads, developer platform owners, and procurement teams buying multi agent coding products should treat maximum concurrent agents and orchestration architecture as first class requirements. Security reviewers should examine shared workspace permissions because workers self assign tasks. Benchmark owners may need organization level suites beyond single agent ProgramBench slices.

What to do next

Pilot Agensh style decentralized loops on one high latency internal workflow this quarter. Measure end to end pass rate and wall clock time at 1, 16, and 128 workers before committing spend on a central orchestrator upgrade.

What to watch

Watch for independent replications on non Microsoft harnesses and open models. Track whether vendors ship production grade shared workspace isolation and audit logs, not only demo scale agent counts.

Sources

  1. Primary. arXiv, Agensh: Scaling Organizational Intelligence to 1,024 Agents (23 September 2026). Defines the harness, ProgramBench scaling curves, and 1,024 agent pandoc result.
  2. Secondary. arXiv recent list, Computer Science submissions for 23 September 2026 (23 September 2026). Confirms submission date and multi agent systems subject tag.

Newsletter

Get the dispatch

One field. One email when we publish. Privacy.