Interview · 2 min read

Engineering Is Not Insight: Arvind Narayanan on CRUX Shadow Evaluations and the Agent Research Gap

Compiled from public record: Princeton's Arvind Narayanan and CRUX colleagues find frontier agents can engineer research pipelines but fail open-ended scientific judgment.

By Classy AI News · August 1, 2026

Engineering Is Not Insight: Arvind Narayanan on CRUX Shadow Evaluations and the Agent Research Gap

Forecasts of explosive AI progress often assume agents will soon automate AI research itself. A July 30 paper from the CRUX collaboration offers a sobering counterpoint—and Arvind Narayanan, Princeton computer scientist and core team member, frames the stakes plainly in the published record.

"Our results provide early evidence that today's frontier models cannot solve weeks-long, open-ended AI research questions."

This reconstruction draws on Narayanan's co-authored paper, Can AI agents conduct open-ended AI research? Early evidence from two case studies, and materials at cruxevals.com. No private interview was conducted.

Why shadow evaluations exist

Narayanan and colleagues—including Sayash Kapoor, Andrew Schwartz, Stephan Rabanser, and researchers from the UK AI Security Institute, Georgetown CSET, and Stanford—argue that existing evaluations miss the hardest part of research. Narrow benchmarks test verifiable metrics; blind peer review of AI-generated papers is stochastic and noisy.

Their alternative: shadow evaluations. Take the central research question from a high-quality unpublished paper, give a frontier agent six days and thousands of dollars in compute, and ask the original authors to grade the output as they would a conference submission.

What the agents were given

The team partnered with authors of two unpublished NeurIPS 2026 submissions. Agents received the research question, six days of wall-clock time, $3,000 in Anthropic API credits, GPU credits, and full VM plus open-web access. The goal: produce a paper worthy of a top-tier conference.

The best-performing setup used OpenClaw with Claude Opus 4.8 after dry runs with models from OpenAI and Anthropic. A robustness check repeated one paper with GPT-5.6 Sol and Codex under the same budgets.

Code snippets on a folder represent the executable engineering agents completed without human help

The verdict: engineering yes, research no

Both papers were unambiguous rejections. On Paper 1 (Personas), the agent scored 2/6 overall; on Paper 2 (TabPFN), 1/6. Agents completed engineering without human help yet failed to make substantial progress on the research questions.

Narayanan's team identified five recurring failure modes: poor judgment about publishability bars; uncreative responses to design shortcomings; ineffective backtracking from dead ends; poor resource awareness—with less than 50% of API budget spent; and instruction drift on exploration time and paper length.

The robustness check matters

Repeating one evaluation with GPT-5.6 Sol and Codex reproduced nearly every failure mode. Narayanan's conclusion is explicit—the results are not simply artifacts of a deficient scaffold.

Python code on a transparent screen reflects the engineering tasks agents handled successfully

What Narayanan is not claiming

Shadow evaluations involve non-blinded author reviews, which may introduce bias. The generated papers were unambiguously poor—but CRUX released artifacts for independent expert judgment. Follow-up work with GPT-5.6 Sol, Opus 5, and Fable 5 is planned.

The policy implication

If agents can execute code pipelines but cannot judge when a research direction has failed, the path to recursive self-improvement runs through evaluation design—not just bigger models.

An interactive whiteboard with coded data evokes the presentation bar agents failed to meet

Sources

Newsletter

Get the dispatch

One field. One email when we publish. Privacy.