The Harness Question: Chollet on What ARC-AGI-3 Actually Measures
ARC Prize co-founder François Chollet draws a line between fair API settings and custom harnesses — and says OpenAI's retained-reasoning run is in bounds if costs are reported.
When OpenAI reported that GPT-5.6 Sol reached 38.3 percent on ARC-AGI-3 after enabling retained reasoning and compaction in its Responses API, the number landed like a challenge flag on a field that had just watched Claude Opus 5 post a verified 30.16 percent under ARC Prize's standardized harness.
François Chollet, co-founder of ARC Prize, has addressed that question directly in public posts this week. This article reconstructs his documented statements — not a private interview.
Two kinds of test setups
Chollet separated harnesses custom-made to solve the benchmark from general-purpose API settings available to all API users, which he described as fair game per The Decoder on July 30, 2026.
Under ARC Prize's standardized setup, GPT-5.6 Sol scored 7.78 percent on the semi-private ARC-AGI-3 evaluation at maximum reasoning effort. OpenAI's 38.3 percent used its Responses API with retained reasoning and compaction.
Parity and cost
Chollet noted extensive back-and-forth with OpenAI about compaction testing. Different provider settings create a parity issue he considers acceptable as long as settings and cost are clearly reported.
OpenAI said retained reasoning consumed roughly six times fewer output tokens while tripling the public-set score from 13.3 percent to 38.3 percent.
Verified versus experimental scores
Claude Opus 5 holds the highest verified ARC-AGI-3 score at 30.16 percent. GPT-5.6 Sol's verified semi-private score is 7.78 percent. OpenAI's 38.3 percent is an orchestration result on public tasks.
Chollet's public criteria — no custom benchmark knowledge, general API settings allowed, costs disclosed — offer a pragmatic middle path for an industry still negotiating how to score reasoning models.
Sources
- The Decoder — OpenAI claims GPT-5.6 Sol beats Opus 5 on ARC-AGI-3 with its latest API and two additional settings (July 30, 2026)
- OpenAI — How enabling two settings tripled our scores on the ARC-AGI-3 benchmark (July 29, 2026)
- ARC Prize — GPT-5.6 Sol — ARC-AGI Results (July 9, 2026)
- ARC Prize — Claude Opus 5 — ARC-AGI Results (July 24, 2026)