Two Scoreboards: What OpenAI's ARC-AGI-3 Orchestration Run Actually Proves
OpenAI's 38.3% ARC-AGI-3 figure and Opus 5's verified 30.16% answer different questions — and the gap defines how enterprise buyers should read benchmark wars in 2026.
OpenAI's July 29 post on ARC-AGI-3 is not a new model launch. It is an argument about what benchmark scores measure when inference-time memory becomes a product feature.
The company reported that enabling retained reasoning and compaction in its Responses API raised GPT-5.6 Sol's public-set score from 13.3 percent to 38.3 percent while using roughly six times fewer output tokens per game. That figure exceeds Claude Opus 5's verified 30.16 percent on the standardized ARC Prize harness — but it was not produced under that harness.
Two scoreboards, one headline
Under ARC Prize's official evaluation, GPT-5.6 Sol scored 7.78 percent on the semi-private ARC-AGI-3 set at maximum reasoning effort. OpenAI's higher number reflects production settings: carrying reasoning across turns and compacting older context.
Both results can be true. They answer different questions.
Why labs care about harness parity
François Chollet said general-purpose API settings are fair game if costs are disclosed. He also acknowledged that standardized ARC runs may have disadvantaged OpenAI relative to APIs that already preserved reasoning between steps.
The commercial subtext
Token efficiency matters as much as peak score. OpenAI's sixfold token reduction while tripling performance is an inference-economics story for Codex and ChatGPT deployments.
Anthropic's verified Opus 5 lead remains the apples-to-apples leaderboard result. OpenAI's 38.3 percent is an orchestration experiment on public tasks.
What buyers should track
Demand three numbers: verified harness score, production-configuration score, and token or dollar cost per task.
Sources
- OpenAI — How enabling two settings tripled our scores on the ARC-AGI-3 benchmark (July 29, 2026)
- The Decoder — OpenAI claims GPT-5.6 Sol beats Opus 5 on ARC-AGI-3 with its latest API and two additional settings (July 30, 2026)
- ARC Prize — GPT-5.6 Sol — ARC-AGI Results (July 9, 2026)
- ARC Prize — Claude Opus 5 — ARC-AGI Results (July 24, 2026)