Biotech · 4 min read

The Closed Loop: Why AI Drug Discovery Needs Negative Data and FAIR Lab Infrastructure

MIT Technology Review and Cytiva's Paul Belcher argue AI drug discovery's next bottleneck is closing the data loop — especially negative results and FAIR lab infrastructure.

By Classy AI News · July 29, 2026

The Closed Loop: Why AI Drug Discovery Needs Negative Data and FAIR Lab Infrastructure

Drug discovery's AI story has shifted. The question is no longer whether models can generate molecules — it is whether the data flowing back from the lab is good enough, structured enough, and honest enough to make those models smarter over time.

A July 27 MIT Technology Review piece, produced in partnership with life-sciences company Cytiva, frames the next bottleneck as the data loop between computational design and wet-lab validation. The argument deserves attention because it comes at the moment when AI-originated drug candidates are entering clinical trials — but none has yet received full FDA approval.

From hit identification to predictive design

Paul Belcher, Cytiva's director of protein research strategy, describes a shift from empirical screening to predictive design. Instead of physically testing compound libraries at scale, companies increasingly use AI to design candidates from scratch and predict target interactions before committing anything to R&D.

"AI does away with that," Belcher told MIT Technology Review, referring to the old constraint of physical library size. "And it can help eliminate low-quality candidates before you have to physically test them, saving time and resources."

What AI cannot yet do reliably, Belcher added, is predict kinetics or developability. Every AI-generated candidate still needs wet-lab validation — and the volume of candidates is rising faster than lab infrastructure built for binary yes-or-no screening can handle.

Research scientist synthesizing compounds in a modern laboratory

The data wall

Many early AI models trained on publicly available datasets are hitting what Belcher calls a data wall. Because models share access to the same public data, they converge on similar conclusions with diminishing returns.

Publication bias compounds the problem. "Most publicly available datasets and scientific publications focus exclusively on positive results," Belcher said. "No one wants to share their failures. This bias is almost like having one hand tied behind your back. AI models can identify patterns associated with success, but they lack the comprehensive understanding of failures that would make predictions more reliable."

The negative data — failed experiments, compounds that do not bind — remains buried in lab notebooks rather than feeding back into training pipelines. Belcher joked that there should be "a journal of negative data."

Integrity risks in an AI training era

Fabrication has also become easier. Belcher cited research by microbiologist Elisabeth Bik finding that almost 4 percent of biomedical papers contained duplicated or manipulated images — a figure from 2016, before generative AI made fabrication trivial.

"Manipulated or faked data has always been a problem in science, but in the AI world, especially when used to train models, it could have potentially disastrous consequences," Belcher said.

Cytiva pointed to its Image Integrity Checker, which uses secure hash algorithms to detect tampered scientific images — technology publishing houses are beginning to adopt as a verification standard.

Scientist working in a chemical laboratory with test equipment

Labs-in-the-loop as the operating model

Belcher's vision for the next phase is fully autonomous labs — "dark labs" or labs-in-the-loop — cycling through prediction, testing, and optimization around the clock, feeding results back into AI models to guide the next experimental round.

That requires integration most labs lack today. "You can have the best technology in the world, but if it's a closed ecosystem — if the user can't get the data out — it doesn't do any good," Belcher noted.

An integrated infrastructure generating FAIR data — findable, accessible, interoperable, and reusable — would close the loop between dry-lab computation and wet-lab execution. Better starting points plus more optimization rounds should, in theory, produce better clinical candidates with fewer liabilities.

The clinical proof gap

MIT Technology Review noted that no drug discovered primarily through AI-driven design has yet received full FDA approval, though Belcher expects that to change within two to three years.

Industry counts cited elsewhere in July 2026 coverage place roughly 170 AI-originated programs in clinical development, with Insilico Medicine's Rentosertib representing the first generative-AI-designed candidate to complete a peer-reviewed Phase IIa trial. The industry's consensus at the June 2026 BIO convention quietly shifted from "can AI make drugs?" to "do AI's drugs actually work in the clinic?"

The compute-versus-clinic tension

Belcher acknowledged the cost tension. Stanford research cited in the piece found frontier AI training costs more than doubling every year since 2016 — adding pressure to a sector where clinical development already runs $1 billion to $2.5 billion per approved drug.

Belcher remains optimistic about balance: "As long as the cost of compute doesn't ever outweigh the cost of clinical development, I think AI is going to be an advantage."

The verified takeaway for July 2026 is structural, not hype-driven: AI drug discovery's next gains depend less on better generative models and more on closing the data loop — negative results, FAIR infrastructure, and labs that can feed learning back into design at scale.

Sources

Newsletter

Get the dispatch

One field. One email when we publish. Privacy.