Analysis · 2 min read

Two Scoreboards: What OpenAI's ARC-AGI-3 Orchestration Run Actually Proves

OpenAI's 38.3% ARC-AGI-3 figure and Opus 5's verified 30.16% answer different questions — and the gap defines how enterprise buyers should read benchmark wars in 2026.

By Classy AI News · July 30, 2026

Two Scoreboards: What OpenAI's ARC-AGI-3 Orchestration Run Actually Proves

OpenAI's July 29 post on ARC-AGI-3 is not a new model launch. It is an argument about what benchmark scores measure when inference-time memory becomes a product feature.

The company reported that enabling retained reasoning and compaction in its Responses API raised GPT-5.6 Sol's public-set score from 13.3 percent to 38.3 percent while using roughly six times fewer output tokens per game. That figure exceeds Claude Opus 5's verified 30.16 percent on the standardized ARC Prize harness — but it was not produced under that harness.

Vibe coding concept with code snippets and diagrams

Two scoreboards, one headline

Under ARC Prize's official evaluation, GPT-5.6 Sol scored 7.78 percent on the semi-private ARC-AGI-3 set at maximum reasoning effort. OpenAI's higher number reflects production settings: carrying reasoning across turns and compacting older context.

Both results can be true. They answer different questions.

Why labs care about harness parity

François Chollet said general-purpose API settings are fair game if costs are disclosed. He also acknowledged that standardized ARC runs may have disadvantaged OpenAI relative to APIs that already preserved reasoning between steps.

Business people communication in a corporate office setting

The commercial subtext

Token efficiency matters as much as peak score. OpenAI's sixfold token reduction while tripling performance is an inference-economics story for Codex and ChatGPT deployments.

Anthropic's verified Opus 5 lead remains the apples-to-apples leaderboard result. OpenAI's 38.3 percent is an orchestration experiment on public tasks.

Dimly lit office cubicle with a glowing computer screen

What buyers should track

Demand three numbers: verified harness score, production-configuration score, and token or dollar cost per task.

Dark office cubicle with an illuminated computer screen

Sources

Newsletter

Get the dispatch

One field. One email when we publish. Privacy.