Interview · 2 min read

The Harness Question: Chollet on What ARC-AGI-3 Actually Measures

ARC Prize co-founder François Chollet draws a line between fair API settings and custom harnesses — and says OpenAI's retained-reasoning run is in bounds if costs are reported.

By Classy AI News · July 30, 2026

The Harness Question: Chollet on What ARC-AGI-3 Actually Measures

When OpenAI reported that GPT-5.6 Sol reached 38.3 percent on ARC-AGI-3 after enabling retained reasoning and compaction in its Responses API, the number landed like a challenge flag on a field that had just watched Claude Opus 5 post a verified 30.16 percent under ARC Prize's standardized harness.

François Chollet, co-founder of ARC Prize, has addressed that question directly in public posts this week. This article reconstructs his documented statements — not a private interview.

Giant chess pieces in a surreal maze representing strategic benchmark design

Two kinds of test setups

Chollet separated harnesses custom-made to solve the benchmark from general-purpose API settings available to all API users, which he described as fair game per The Decoder on July 30, 2026.

Under ARC Prize's standardized setup, GPT-5.6 Sol scored 7.78 percent on the semi-private ARC-AGI-3 evaluation at maximum reasoning effort. OpenAI's 38.3 percent used its Responses API with retained reasoning and compaction.

Parity and cost

Chollet noted extensive back-and-forth with OpenAI about compaction testing. Different provider settings create a parity issue he considers acceptable as long as settings and cost are clearly reported.

OpenAI said retained reasoning consumed roughly six times fewer output tokens while tripling the public-set score from 13.3 percent to 38.3 percent.

Human and robotic hands reaching toward a quantum energy symbol

Verified versus experimental scores

Claude Opus 5 holds the highest verified ARC-AGI-3 score at 30.16 percent. GPT-5.6 Sol's verified semi-private score is 7.78 percent. OpenAI's 38.3 percent is an orchestration result on public tasks.

Abstract blue and white layered shapes with soft lighting

Chollet's public criteria — no custom benchmark knowledge, general API settings allowed, costs disclosed — offer a pragmatic middle path for an industry still negotiating how to score reasoning models.

Sources

Newsletter

Get the dispatch

One field. One email when we publish. Privacy.