The Reconstruct Bottleneck: Anthropic and Andon Labs Put Frontier Models in the Cockpit of a $129 Drone
Anthropic's Project Pilot and Andon Labs' new Drone-Bench benchmark test whether frontier AI models can autonomously locate and follow a person with a consumer drone — and reveal a six-month gap between peak capability and reliable performance.
When Software Agents Learn to Fly
On July 24, 2026, Anthropic's Frontier Red Team published Project Pilot, a research program asking a deceptively simple question: can today's frontier AI models autonomously control a drone to find and follow a specific person indoors?
The answer, according to Anthropic and its evaluation partner Andon Labs, is partially — and that partial answer may be more consequential than a clean yes. Working with a DJI Tello EDU quad-rotor that currently retails for $129, the teams built a locate-and-follow surveillance task inside an office environment and decomposed it into five benchmarkable sub-tasks. The result is Drone-Bench, a new evaluation measuring how well AI models write the code that turns cheap, off-the-shelf hardware into an autonomous aerial tracker.
This is not a toy demo dressed up as research. Anthropic frames Project Pilot as the aerial counterpart to earlier embodied-AI work — Project Vend, in which models ran a small shop, and Project Fetch, which explored robots as intermediaries between digital intelligence and physical objects. As the lab noted in its July 24 post, model capability for using off-the-shelf robots is approaching the ease with which coding agents use software tools. Drones add a particular urgency: they are widely available to professionals and hobbyists alike, and they carry dual-use implications from agricultural monitoring to military targeting.
Five Tasks, One Chain
Drone-Bench breaks the surveillance mission into five sequential capabilities, each scored against a human–AI team baseline that Andon Labs built using coding agents:
- Reconstruct — Turn office videos into a 3D model and produce a function that slices it into a 2D obstacle map.
- Localize — Given video frames with known poses, match the drone's current camera view to its position on that map.
- Navigate — Plan and fly a collision-free path between rooms, continuously calling Localize to correct for noisy controls.
- Detect — Build a face-based detector from a reference photo and return bounding boxes for the target person in each frame.
- Follow — Use those bounding boxes to keep the target centered and at a stable distance as they move.
In the benchmark, each task runs in isolation with clean upstream artifacts provided — a design choice that prevents a weak Reconstruct from unfairly tanking Navigate scores, but also means per-task success does not automatically translate to end-to-end reliability. Andon Labs ran 10 simulations per model across 15 frontier models from three developers, including GPT-4o, o1, o3, Claude Opus 4 through Opus 4.8, Fable 5, GPT-5 through GPT-5.6 Sol, and Gemini 2.5 Pro through Gemini 3.1 Pro.
The overall trend is unmistakable: newer models progress further on every sub-task. Detection and following are where models perform best; reconstruction and localization remain the hard ceiling.
Fable 5 Leads — Until It Hits a Wall
The top performer was Claude Fable 5, which Anthropic reports surpassed the human–AI baseline on four of five tasks, falling short only on Reconstruct. On the real drone, Fable 5 detected and followed the reference person more closely than the baseline algorithm.
But end-to-end autonomy told a different story. Reconstruction errors compounded through localization and navigation, and Fable 5 was unable to autonomously navigate between rooms. Anthropic published video of the model confidently flying toward what it believed was a doorway — and striking a wall.
Andon Labs' data sharpens the picture. No model has yet beaten the Reconstruct baseline (scored at 82.2%). Because end-to-end success requires beating the baseline on all five tasks in sequence, end-to-end success remains at zero percent across every model tested. The best frontier model clears four tasks on a good run, but a typical run strings them together only about 6% of the time.
There is also a consistency gap that should worry anyone planning deployment timelines. Andon Labs estimates the average submission trails the best submission by roughly six months: what the frontier achieves once in a peak run today, a typical run may not match until half a year later. Fable 5, for instance, beats the baseline on just 2% of first submissions but 52% of best submissions within a 10-attempt run — the largest first-to-best jump of any model in the eval.
Anthropic highlighted encouraging signs in Fable 5's reasoning process. In one submission, the model estimated the drone camera's tilt to within four degrees of the true value by analyzing grout lines on the floor and extrapolating to a vanishing point. In another, it built a 2D top-down reconstruction of the Follow environment to test its code locally before submitting — catching bugs before burning an evaluation attempt.
The Oversight Clock Is Ticking
Project Pilot's policy framing is as important as its benchmark numbers. Anthropic explicitly chose a surveillance-style locate-and-follow task because it mirrors capabilities with legitimate uses — search and rescue, disaster response, lawful public safety — while remaining susceptible to abuse through overreach or unaccountable private deployment.
The lab draws a parallel to agentic coding: early deployments required human approval for nearly every tool call; within months, models gained trust to execute long-horizon tasks with minimal intervention. Hardware control, Anthropic argues, will follow the same arc. At low capability and reliability, human oversight saves time and prevents costly mistakes. Once models reliably pass a realistic human–AI baseline, pressure to treat oversight as a cost rather than a safeguard intensifies — especially in domains touching physical security and individual privacy.
Andon Labs reinforces that governance cannot remain a lab-only conversation. "Whether and how we handle that is a question for the public, not just AI labs," the team wrote in its Drone-Bench release. The benchmark is designed so that no AI lab can train on the eval — preserving its value as an independent capability signal.
What the Numbers Do — and Do Not — Prove
Both Anthropic and Andon Labs are careful about limitations. The drones moved at slow speeds. Testing occurred in one office floorplan with a limited number of people, all consented members of the experiment team. Outdoor crowd scenarios were not evaluated. These constraints mean Project Pilot provides a directional signal, not an operational readiness assessment.
Still, the signal is clear enough to act on. Frontier models are approaching — but have not yet crossed — the threshold where they could autonomously recreate a basic indoor surveillance demo at least as capable as a human expert working with modern coding tools. Reconstruct is the remaining bottleneck; once a model clears it, end-to-end performance could jump discontinuously even though progress on individual sub-tasks has been gradual.
Andon Labs projects that the next frontier model's best solution may pass all five tasks, with typical runs catching up within another six months. That timeline, if it holds, would compress the window for policymakers, civil-society groups, and industry to agree on norms for AI-controlled aerial systems — before capability outruns the deliberation.
For robotics watchers, Project Pilot also reframes a familiar debate. The question is no longer whether language models can talk about drones. The question is whether they can reliably build the software stack that makes one fly, see, and follow — and how far behind the headline best-case run the average deployment will lag.
Sources
- Anthropic — Project Pilot: Can AI control a drone? (July 24, 2026)
- Andon Labs — Drone-Bench (July 2026)
- Anthropic — Project Fetch: Phase two (prior embodied-AI context cited in Project Pilot)