Harbor Index Condenses 54 Agent Benchmarks Into 82 High Signal Tasks
Harbor Index 1.0 selects 82 difficult tasks from more than 6,600 candidates across 54 benchmarks, capping the best model harness pass rate at 28 percent and cutting eval cost to a few hundred dollars per run.
What changed
On 3 September 2026, researchers posted Harbor Adapters and Harbor Index to arXiv (2609.04298) and released Harbor Index 1.0 at harbor-index.org. The project ports more than 80 agentic benchmarks through unified adapters, runs eight models across 54 benchmarks, and distills a curated meta dataset of 82 tasks spanning 29 source benchmarks and seven domains.
The authors report that no evaluated model harness configuration exceeds a 30 percent pass rate on the index. The strongest configuration, GPT-5.5 with Codex, reaches 28.0 percent. Building the full adapted suite consumed roughly 226 billion tokens and more than $300,000 in compute, while a Harbor Index run costs a frontier agent only a few hundred dollars according to the project site.
Why it matters
Agent evaluation has fractured into dozens of incompatible environments, making vendor claims hard to compare and expensive to reproduce. Harbor Index answers a procurement problem: teams need a compact benchmark that still punishes frontier models. If the best public configuration stops below 30 percent, leaderboard chasing on easy tasks loses decision value.
The funnel matters as much as the score. Candidates pass a difficulty filter, automated AI audit, human audit, and repeated fix loops before entering the 82 task set. That process is designed to preserve challenge and domain breadth while removing stochastic noise, aligning with the authors' cited work arguing teams do not need to run every eval.
Who is affected
Applied AI leads choosing eval suites for coding agents, security tools, and long horizon research assistants.
Model vendors competing on agentic claims who must now defend performance on an audited hard subset, not cherry picked wins.
Safety and red team groups who can reuse Harbor adapters to port cyber and software engineering benchmarks under one harness.
What to do next
Pilot Harbor Index on your top two agent stacks before renewing eval contracts. Compare pass rates and failure modes against your internal task library to see whether gaps are harness induced or domain specific.
What to watch
Whether major labs adopt Harbor Index as a standard reference in system cards, and whether adapter parity experiments hold when new native harnesses ship with GPT-6 class models.
Sources
- Primary. arXiv, Harbor Adapters and Harbor Index: Infrastructure and a Curated Meta Dataset for Large Scale Agentic Evaluation (3 September 2026). Adapter scope, model sweep, and 28.0 percent ceiling.
- Primary. Harbor Index, Introducing Harbor Index 1.0 (3 September 2026). Funnel stages, 82 task composition, and cost comparison to full suite runs.
- Secondary. arXiv, You don't need to run every eval (2026). Theoretical basis for compact high signal evaluation cited by Harbor Index authors.