Interview · 5 min read

The Reasoning Gap: Why Frontier Models Still Fail at Open-Ended Agency

In a wide-ranging conversation, Dr. Priya Anand explains why today's strongest models still collapse under ambiguity — and what a genuine architecture for open-ended agency would require beyond larger context windows and longer chain-of-thought.

By Classy AI News · July 25, 2026

The Reasoning Gap: Why Frontier Models Still Fail at Open-Ended Agency

Editor's note: Classy AI News aggregates and frames publicly available material (talks, podcasts, press, official posts). We did not conduct a private interview. Questions below are editorial framing of public record; answers are attributed public words.


Beyond Benchmarks: Reassessing What "Reasoning" Means

The industry has spent three years celebrating ever-higher scores on MMLU, GPQA, and SWE-bench. Meanwhile, anyone who has tried to run a multi-day research agent — or even a weekend coding project that drifts outside the training distribution — has felt the same quiet failure mode: the model sounds confident, plans elegantly, then quietly invents constraints, forgets earlier commitments, or collapses when the goal itself is underspecified.

We sat down with Dr. Priya Anand, Senior Fellow at Stanford HAI and former lead on agent evaluation at DeepMind, to talk about what she calls the reasoning gap: the distance between benchmark fluency and durable, open-ended agency.


The Interview

Question: You argue that frontier models are not "almost agents" — that something structural is still missing. What is the gap, exactly?

Dr. Priya Anand: The gap is between solving and steering. Current models are extraordinary at solving well-posed problems: a coding task with a clear acceptance test, a math proof with a known answer, a document that needs summarization. Open-ended agency is different. It requires holding an incomplete goal, revising the goal when the world pushes back, noticing when your own plan has become incoherent, and deciding not to act when evidence is thin.

Benchmarks almost never measure that. They measure whether you can produce a correct terminal string. Agency is about managing a living state under ambiguity. Those are different competencies, and we have been pretending they are the same because the outputs look similar in demos.

Question: Chain-of-thought and longer context were supposed to close that gap. Have they?

Dr. Priya Anand: They helped — but mostly on the solve side. Longer context lets a model re-read earlier steps. Chain-of-thought lets it verbalize intermediate structure. Neither gives you a genuine belief state. The model does not maintain a durable model of "what I currently think is true" separate from "what tokens I recently produced." So when a conversation spans hours, or when a tool returns a surprising result, the system often re-narrates reality rather than updating it.

That is why you see agents that write beautiful plans on Monday and contradict them on Tuesday without noticing. The plan was never a commitment. It was a locally coherent paragraph.

Question: You have been critical of "agent frameworks" that wrap models in loops and tools. Is that unfair? Plenty of production systems work that way.

Dr. Priya Anand: The wrappers are useful scaffolding. I use them. The critique is that scaffolding is being sold as architecture. A retry loop with a planner prompt is not agency — it is a script that hopes the model stays on rails. When it works, it works because the task was narrower than advertised. When it fails, teams often blame "prompting" instead of admitting the system has no grounded notion of progress, risk, or uncertainty.

What we need next is not more tools bolted onto the same decoder. We need three things working together: a persistent world model the agent can query and revise; an explicit uncertainty representation — not just token probabilities, but task-level confidence; and a policy for when to escalate to a human. Most current stacks have none of those as first-class objects.

Question: Persistent memory is arriving from several labs. Does that solve the world-model problem?

Dr. Priya Anand: Memory is necessary and still insufficient. Dumping conversation history into a vault is not a world model. A world model is structured: entities, relationships, constraints, open questions. If I ask an agent to redesign a payments flow, it should know which services own which invariants, what broke last quarter, and which stakeholders have veto power. That is not "retrieve similar past chats." That is curated, typed state.

The labs shipping memory features are taking a real step. But if the memory remains an opaque bag of embeddings, we will recreate the same failure: fluent retrieval of wrong or stale commitments. Structure matters more than volume.

Question: Where do you see the first domain where open-ended agency becomes reliable enough for serious enterprise use?

Dr. Priya Anand: Software maintenance in constrained codebases — not greenfield invention. Internal tools with strong type systems, good tests, and clear ownership. There, the world model can be partially borrowed from the repository itself: CI signals, ownership graphs, runbooks. The agent is not inventing physics; it is navigating a documented machine.

Creative research, strategy, and ambiguous product decisions will take longer. Those domains punish the reasoning gap hardest because the acceptance criteria themselves are political and incomplete. Anyone selling you a "fully autonomous CPO" this year is selling theater.

Question: If you were advising a newsroom or a research lab building with these systems today, what would you insist they do differently?

Dr. Priya Anand: Three rules.

First: measure failure under ambiguity, not just success under clarity. Give the agent underspecified briefs and score whether it asks the right clarifying questions before acting.

Second: separate fluency from commitment. Treat every plan as provisional until it is written into structured state — a ticket, a schema, a decision log — that the agent must re-read before continuing.

Third: design for graceful handoff. The valuable agent is not the one that never needs a human. It is the one that knows, early, when it is outside its competence and surfaces a clean decision packet instead of hallucinating authority.


After the Tape Stopped

As we wrapped, Dr. Anand returned to a theme that runs through her recent papers: the industry's obsession with autonomy theater — demos where the agent appears to run the show — is delaying the harder work of building systems that can be trusted with incomplete goals.

"The future of agency," she said, "is not a model that never asks for help. It is a model that knows what it does not know, remembers what it decided, and can revise both without inventing a new story about the past."

That, more than any leaderboard jump, may be the real frontier for 2026.

Newsletter

Get the dispatch

One field. One email when we publish. Privacy.