The Closed Loop: Why AI Drug Discovery Lives or Dies on Lab Integration, Not Model Size
Cytiva’s Paul Belcher told MIT Technology Review that AI drug discovery succeeds only when autonomous labs feed prediction and validation back into models — and when negative data stops staying in notebooks.
Drug discovery's AI story is no longer about whether models can predict binding. It is about whether predictions ever return to the lab in a loop tight enough to change what enters clinical trials.
A July 27, 2026 MIT Technology Review feature — produced in partnership with Cytiva — maps where that loop breaks and where it is finally closing.
From screening libraries to designing candidates
Paul Belcher, Cytiva's director of protein research strategy, describes a shift from empirical screening to predictive design: instead of physically testing millions of compounds, companies use AI to design candidates from scratch and model target interactions before committing R&D spend.
Hit identification — finding molecules that bind a disease-related protein — remains one of the most promising early-stage applications. Belcher notes AI can eliminate low-quality candidates before wet-lab testing, but every AI-generated hit still requires laboratory validation because models cannot yet reliably predict kinetics or developability.
The data wall nobody advertises
Many early models trained on public datasets are hitting what Belcher calls a data wall: shared training corpora produce diminishing returns, and those datasets were not built with AI-native structure, labeling, or diversity in mind.
Publication bias makes it worse. "Most publicly available datasets focus exclusively on positive results," Belcher said. "No one wants to share their failures." Without negative data, models learn success patterns but lack the comprehensive failure signal that would make predictions more reliable.
Fabrication risk has compounded the problem. Belcher cited Elisabeth Bik's finding that nearly 4 percent of biomedical papers contained duplicated or manipulated images — a rate from 2016, before generative AI made fabrication trivial.
Labs-in-the-loop as the operating model
Belcher's forward view is autonomous dark labs: facilities cycling through prediction, testing, and optimization around the clock, feeding results back into models to guide the next experimental round.
That vision depends on integration — interoperable instruments, FAIR data at scale, and workflows where information flows in both directions. "You can have the best technology in the world," Belcher noted, "but if it's a closed ecosystem — if the user can't get the data out — it doesn't do any good."
What approval would prove
No drug discovered primarily through AI-driven design has received full FDA approval yet. Belcher expects that to change in the next two to three years — a milestone that would validate AI as a discovery tool without by itself proving improved clinical success rates.
The holy grail remains full in silico efficacy and toxicity prediction. Belcher is optimistic about a balance between AI and wet work — as long as compute costs never exceed the cost of clinical development.
For 2026, the biotech AI transition is operational: from pilot projects to integrated systems where the lab and the model share a single feedback cycle — or neither delivers on the promise.
### Sources
- MIT Technology Review — Closing the data loop in AI-driven drug discovery (July 27, 2026)
- Drug Discovery News — The 2026 AI power shift (2026)