Benchmark Literacy: Why ARC-AGI-3 Needs Labels, Not Just Leaders
The OpenAI–ARC Prize dispute is not about cheating — it is about whether buyers will finally demand labeled scores, token costs, and harness types before treating benchmarks as AGI proxies.
The ARC-AGI-3 exchange between OpenAI and ARC Prize is not a scandal. It is the first honest public fight about a problem the industry has avoided: once models ship with memory, compaction, and tool loops, a single benchmark percentage stops meaning one thing.
OpenAI's July 29 post is clearest when read as economics, not ego. Retained reasoning and compaction raised GPT-5.6 Sol's public-set score from 13.3 percent to 38.3 percent while cutting output tokens per game by roughly sixfold. That is a deployment story — the kind of gain customers actually pay for — even when the verified semi-private harness still shows 7.78 percent because it deliberately amputates memory between actions.
François Chollet's public response, as reported July 30, is equally pragmatic. General-purpose API settings are fair game if costs are disclosed; custom benchmark harnesses are not. That is a reasonable line — but it also admits the old contract is breaking. ARC Prize built a standardized harness to compare weights. OpenAI built a Responses API to compare products. Both are doing their jobs.
The danger is not that labs game benchmarks. The danger is that buyers treat any headline number as a proxy for AGI proximity. Claude Opus 5's verified 30.16 percent and OpenAI's 38.3 percent orchestration result can coexist without either one lying. They measure different layers of the stack.
What should change is disclosure, not outrage. Every frontier score should ship with a label: verified harness, production configuration, tokens consumed, dollars estimated, and whether the result is on public or semi-private tasks. Benchmark literacy is becoming as important as benchmark performance.
Until evaluators publish those fields by default, the winner of each benchmark news cycle will be whoever controls the headline — not whoever controls the cost curve. That is a bad equilibrium for buyers, regulators, and the labs themselves.
The ARC fight is healthy if it forces that transparency. The industry should treat it as a template, not an exception.
Sources
- OpenAI — How enabling two settings tripled our scores on the ARC-AGI-3 benchmark (July 29, 2026)
- The Decoder — OpenAI claims GPT-5.6 Sol beats Opus 5 on ARC-AGI-3 with its latest API and two additional settings (July 30, 2026)
- ARC Prize — GPT-5.6 Sol — ARC-AGI Results (July 9, 2026)
- ARC Prize — Claude Opus 5 — ARC-AGI Results (July 24, 2026)