Research · 3 min read

Claude Pushes a Riemann Zeta Bound to 67.2% — With Humans and Lean in the Loop

An unreleased Claude research build raised a classical lower bound on Riemann zeta zeros on the critical line from 41.6% to 67.2%, with human expert review and a Lean formalization — not a proof of the full hypothesis.

By Classy AI News · August 11, 2026

Claude Pushes a Riemann Zeta Bound to 67.2% — With Humans and Lean in the Loop

The headline result

On August 10, 2026, Anthropic published a research note describing an unexpected mathematical advance from an unreleased research version of Claude. While attempting work related to the Riemann hypothesis, the model improved a longstanding lower bound on the fraction of nontrivial zeros of the Riemann zeta function that lie on the critical line — from 41.6% to 67.2%.

That is not a proof of the full hypothesis, which would require showing 100% of nontrivial zeros satisfy the conjecture. It is, however, the largest single improvement Anthropic reported for this particular bound, and it arrived as a byproduct of a failed attack on the million-dollar problem itself.

Researchers working with laboratory equipment

Why the Riemann zeta function matters

The Riemann zeta function encodes information about the distribution of prime numbers. The Riemann hypothesis — unsolved since 1859 — posits that all nontrivial zeros lie on a specific vertical line in the complex plane. Mathematicians have progressively tightened lower bounds on what fraction of zeros are known to obey the conjecture; Anthropic says human researchers had reached 41.6% before Claude's run.

Progress on surrounding bounds matters because many results in analytic number theory assume the hypothesis or related randomness properties of the primes.

How Claude found the bound

According to Anthropic's account, the work unfolded inside Claude Code across two sessions totaling roughly 31 million output tokens.

Jarred Sumner, an Anthropic staff member without a mathematics background, prompted Claude to "take a real stab" at the hypothesis itself. The model first generated and tested 650 ideas, none of which worked. After encouragement to continue, Claude coordinated about 60 subagents over roughly a day and a half. Those agents ran 2,400 shell commands, wrote hundreds of Python scripts, and performed thousands of numerical checks against known zeta zeros.

Anthropic credits two in-house mathematicians — Levent Alpöge and Ralph Furman — with examining the result, along with external experts Brian Conrey and Dan Goldston, who reviewed the paper on short notice. Claude also produced a Lean formalization that passes standard validation tooling, and staff member Eric Easley helped coordinate that effort.

The technical path, in brief: Claude combined recent work by Baluyot, Goldston, Suriajaya, and Turnage-Butterbaugh — which allows Montgomery-style techniques without assuming the hypothesis — with a 2000 paper by Bombieri, surpassing the prior 41.6% constant.

Abstract editorial image evoking discovery and scrutiny

Verification stack

Anthropic describes a layered validation process rather than treating the model output as self-evident:

  • Internal mathematicians studied the proof and its relationship to prior literature.
  • External number theorists reviewed the write-up.
  • Subagents searched for counterexamples and checked 54 arXiv papers to ensure the bound was novel.
  • A formal Lean proof was generated and validated.

That pipeline matters as much as the number 67.2%. Frontier models can now extend expert-level mathematical literature, but the publication still routes through human and formal verification — not raw model assertion.

What this does and does not show

What it shows: Claude can synthesize decades of human progress on a hard analytic number theory frontier and push a quantitative bound materially forward, even when the original objective — proving the Riemann hypothesis — remains out of reach.

What it does not show: A general solution to open problems in pure mathematics, or a replacement for expert review. Anthropic explicitly states it does not expect the specific techniques to yield a full proof of the hypothesis.

Anthropic also notes Claude was initially skeptical of its own finding, possibly reflecting training on the difficulty of major open problems — a reminder that model confidence and mathematical truth remain decoupled.

Broader research context

The result lands in a month when labs are publicly grappling with agentic capability and evaluation integrity elsewhere in the stack. Within mathematics, the interesting research question is shifting from "Can models solve olympiad problems?" to "Can models extend research frontiers with verifiable artifacts?"

Anthropic has published supplementary materials including the paper, an appendix explaining Claude's discovery process, and the Lean formalization linked from its research page.

Sources

Newsletter

Get the dispatch

One field. One email when we publish. Privacy.