Flash Pricing vs. Ultrafast Tokens: Google and OpenAI Are Optimizing Different Agent Bottlenecks
Google's Gemini 3.7 Flash bets on agent unit economics at $0.75/$3.75 per million tokens; OpenAI's GPT-5.6 Sol Ultrafast bets on 14× speed at frontier capability — two different production constraints.
Two speed stories, one agent economy
August 13, 2026 delivered a paired announcement day: Google shipped Gemini 3.7 Flash to general availability while OpenAI previewed GPT-5.6 Sol Ultrafast to an invite-only customer set. Both labs framed the releases around agents and real-time workflows — but the go-to-market shapes diverge sharply.
Google's model is a Flash-tier refinement of 3.6 Flash, available broadly through the Gemini API, Google AI Studio, Gemini Enterprise, and Spark for Pro/Ultra subscribers in more than 160 countries. OpenAI's Ultrafast is the same GPT-5.6 Sol model on a faster inference track powered by Cerebras, currently limited to a small preview cohort.

What Google is optimizing for
Google's blog positions 3.7 Flash as a workhorse for coding and agents, three weeks after 3.6 Flash. Published benchmark deltas versus 3.6 Flash include:
- FrontierCode 1.1 Main: 43.6% vs 34.4%
- DeepSWE v1.1: 65.3% vs 49.0%
- WebDev Arena Elo: 1588 vs 1538
- GDP.pdf: 34.0% vs 22.0%
- AutomationBench: 30.4% vs 17.0%
The stronger commercial claim may be price. Introductory API rates through December 31, 2026 are $0.75 per million input tokens and $3.75 per million output tokens — half the original 3.6 Flash list rate. Starting January 1, 2027, those rates double to $1.50 and $7.50.
Google also updated Gemini Spark, its 24/7 personal agent for Pro/Ultra subscribers, to run on 3.7 Flash starting August 13.
What OpenAI is optimizing for
OpenAI's Ultrafast preview targets 750 output tokens per second, roughly 14× standard GPT-5.6 Sol speed, via Cerebras hardware. The company highlighted incident response, customer support, financial market analysis, and voice stacks as use cases.
TechCrunch noted OpenAI has not published head-to-head benchmark scores for Ultrafast beyond customer quotes. Jane Street AI engineer John Crepezzi said Cerebras' speed "enables different ways of using the models" in OpenAI's announcement; Podium product lead Courtland Lykins called it "invaluable in our voice stack."
The tier is invite-only while capacity grows.

The structural read
These releases are not the same product category.
Gemini 3.7 Flash competes on cost-per-successful-agent-run — cheaper tokens plus better coding and workflow scores at Flash latency. That matters for developers running long-horizon agent loops where token volume dominates bills.
GPT-5.6 Sol Ultrafast competes on time-to-first-useful-token at frontier capability — keeping the most capable model class while approaching real-time interaction speeds. That matters for voice, trading desks, and live incident triage where seconds change UX.
The labs are no longer arguing only about benchmark ceilings. They are arguing about which constraint binds production agents: unit economics or wall-clock responsiveness.
Risks the marketing omits
Google's January 2027 price doubling is not cosmetic — teams building on introductory rates need migration budgets. The New Stack flagged that DeepSWE leaderboard gains may reflect eval-specific tuning; production variance remains an open question.
OpenAI's Cerebras partnership is a capacity rental for speed, not a new model. Ultrafast availability is gated; enterprises cannot plan around it yet.
Both releases arrive as multi-agent safety becomes a live concern — Anthropic's August 13 turf-war research showed independent agents with conflicting orders escalating on shared infrastructure. Faster, cheaper agents multiply that surface area.

Sources
- Google — Introducing Gemini 3.7 Flash (August 13, 2026)
- TechCrunch — OpenAI introduces Ultrafast, a new mode that makes GPT-5.6 Sol work at 14x the speed (August 13, 2026)
- Decrypt — Google and OpenAI Debut Super Fast AI Models (August 13, 2026)
- The New Stack — The AI model that just scored 65% on DeepSWE isn't the one Google promised (August 13, 2026)