Analysis · 4 min read

The World-Model Wager: Is Google Playing a Different AGI Game Than OpenAI and Anthropic?

Analyzing Alberto Romero's hypothesis that Google DeepMind is betting on world models while OpenAI and Anthropic sprint toward recursive self-improvement — with verified Gemini 3.6 Flash benchmark context.

By Classy AI News · July 29, 2026

The World-Model Wager: Is Google Playing a Different AGI Game Than OpenAI and Anthropic?

Alberto Romero's July 28 Substack essay "The Actual Reason Why Google 'Fell Out' of the AI Race Changes Everything" offers a provocative hypothesis: Google DeepMind may be deliberately steering away from the recursive-self-improvement path that OpenAI and Anthropic are betting everything on — not because Google lost, but because CEO Demis Hassabis believes world models, not coding-agent automation, are the route to AGI.

Romero labels this informed speculation. Classy AI News treats it as analysis grounded in public breadcrumbs — not confirmed strategy. What follows separates verified facts from Romero's interpretive frame.

What OpenAI and Anthropic are publicly optimizing for

Romero's essay centers on a divergence in theories of intelligence. OpenAI and Anthropic leadership have repeatedly described AI systems that can autonomously improve AI research — recursive self-improvement (RSI) — as plausible within years, not decades.

Public signals Romero cites include:

  • Engineering staff at both labs reporting that coding agents (Claude Code, Codex) now write most internal code, with humans acting as managers of agent swarms.
  • OpenAI describing GPT-5.6 Sol as having "autonomously post-trained" a successor model.
  • Jack Clark writing in May 2026 that there is a 60% chance RSI is achieved by 2028.
  • OpenAI's March 2026 "code red" restructuring canceling "side quests" to refocus on agents and enterprise after Anthropic's gains.

These are reported behaviors and executive statements, not proof that full RSI has arrived. They do describe a coherent strategic bet: scale compute, automate research loops, move fast.

Where Google looks different on verified benchmarks

Google released Gemini 3.6 Flash on July 21, 2026. Google's official blog frames it as a workhorse model with better coding, knowledge work, and token efficiency — 17% fewer output tokens than Gemini 3.5 Flash on the Artificial Analysis Index, priced at $1.50/$7.50 per million input/output tokens.

Independent testing tells a sharper story about frontier rank. Office Chai reported that Gemini 3.6 Flash scores 50 on Artificial Analysis Intelligence Index v4.1 — identical to Gemini 3.5 Flash — leaving Google's best publicly available model in the same composite tier while Anthropic, OpenAI, and others hold higher-scoring frontier systems.

Ground Truth summarized the independent read: Google shipped a faster, cheaper worker, not a higher composite intelligence score — output throughput roughly 1.84× the predecessor, with cost per task falling, but no leap on the headline index.

Romero's essay interprets that gap as strategic withdrawal from the RSI race. The verified fact is narrower: Google's latest public Flash release did not advance composite frontier intelligence on AA's index, even as it improved efficiency and selected agentic sub-benchmarks.

Monocrystalline silicon chips evoking long-horizon hardware and research bets

Romero's Hassabis hypothesis: world models over RSI

Romero argues Hassabis is betting on world models — systems that simulate physical and social reality — rather than token-prediction agents automating their own codebases. In this reading, OpenAI and Anthropic's RSI trajectory is at best an off-ramp and at worst a dead end; Google's incumbent search and cloud businesses can subsidize a longer research horizon.

Romero explicitly warns there is no official confirmation. Google executives still describe competitive frontier development publicly. The essay asks readers to follow product choices, release cadence, and organizational emphasis — Antigravity, Omni, Gemma-family releases — as clues to a different AGI theory.

Classy AI News cannot verify Hassabis's private convictions. We can note that Google DeepMind has historically invested in robotics, simulation, and multimodal world understanding alongside LLMs — a portfolio consistent with Romero's frame, but also consistent with simply being a large incumbent with multiple bets.

Why the hypothesis matters even if wrong

If Romero is wrong, Google is merely behind on the RSI curve and risking "escape velocity" — Anthropic's 2023 claim that labs training the best 2025–26 models may be uncatchable afterward. If Romero is right, the AI race is not one race but two: a near-term agentic RSI sprint among startups, and a longer world-model arc inside Google that could look like lagging until it does not.

Either way, July 2026 data supports a factual premise: OpenAI and Anthropic are setting the public pace on agents, enterprise API share, and frontier benchmark narratives, while Google's July Flash release optimized efficiency without moving the composite intelligence needle.

Glowing computer chip under microscope representing alternate AGI research paths

What would falsify Romero's read

Google would need to show composite frontier intelligence gains on independent indices, release models clearly competitive with Claude Fable 5 and GPT-5.6 Sol on agentic coding benchmarks, and organize public messaging around automated research loops rather than efficiency-tier Flash updates.

Until then, Romero's essay functions as the most detailed public attempt to explain why Google looks absent from the RSI headlines — as strategic choice rather than mere execution failure.

The bottom line

Verified: Gemini 3.6 Flash launched July 21 with efficiency gains but unchanged AA Intelligence Index composite score. Verified: OpenAI and Anthropic continue public emphasis on coding agents and automated research. Unverified but argued well: Hassabis may be playing a different game entirely.

Readers should hold both layers — fact and hypothesis — at once. The race narrative everyone watches may not be the only race on the board.

Sources

Newsletter

Get the dispatch

One field. One email when we publish. Privacy.