Insilico's DDD Benchmark Asks the Question Pharma Cannot Ignore: Can AI Actually Discover Drugs?
Insilico Medicine's DDD Benchmark as a Service tests frontier models on decontaminated, real-world drug discovery tasks — 300+ evaluations and full-program PCC scenarios.
Beyond exam scores: Insilico launches a drug-discovery benchmark you cannot memorize
Insilico Medicine on August 14, 2026 launched the Drug Discovery and Development (DDD) Benchmark as a Service (BaaS) — a standardized evaluation framework asking whether frontier AI models can make real drug discovery decisions, not just score well on contaminated public tests.
The service, available at dddbench.insilico.com, invites organizations to submit models via standard chat-completions APIs and receive scorecards against expert baselines drawn from Insilico's validated programs.
The contamination problem
AI models have spread rapidly across drug discovery, but a persistent question divides practitioners: are systems finding novel medicines, or passing exams whose answers already sit in training data?
Insilico argues most existing benchmarks suffer from data contamination — inflated scores on memorized test questions that do not predict performance on genuine candidate decisions.
The DDD Benchmark anchors evaluations in decontaminated real-world data and Insilico's own pipeline experience: 31 preclinical candidates nominated in six years, more than ten investigational new drug clearances, and an average timeline to preclinical candidate (PCC) nomination under 18 months versus up to four years in traditional discovery.
Two evaluation suites
Drug Discovery Foundations — more than 300 evaluations built from proprietary out-of-distribution test sets and rigorously decontaminated public data. Tasks span disease biology, molecular property prediction, retrosynthesis, structure-based design, and clinical development.
Drug Candidate Essentials — assesses whether a model can navigate a complete discovery program from initial hit to PCC nomination, with reference baselines anchored in Insilico's validated programs.
Organizations may publish results on a public leaderboard or keep verified score reports internal.
Why agents need a harder yardstick
AI systems increasingly act as agents — planning experiments, reasoning over data, and calling tools via the Model Context Protocol. The DDD Benchmark tests whether those capabilities translate into sound drug discovery decisions rather than strong performance on abstract tasks.
Insilico CEO Alex Zhavoronkov said in the company's announcement:
The rapid progress of AI has made one question more urgent than ever: can these models actually discover drugs? … The DDD Benchmark converts that real-world experience into a rigorous, standardised yardstick that the whole field can use.
The framework builds on Insilico's Pharma.AI platform and MMAI Gym post-training environment for scientific AI.
Pipeline context
Insilico's lead program, Rentosertib (ISM001-055) — an AI-discovered TRAF2/NCK-interacting kinase inhibitor — is in Phase III for idiopathic pulmonary fibrosis. Separately, on August 12, 2026, the company nominated ISM0900, an oral Lp(a) inhibitor, as its seventh preclinical candidate of 2026.
The benchmark does not replace regulatory validation. It offers a common ruler for comparing frontier models before they enter high-stakes discovery workflows.
### Sources
- Labmate Online — Novel benchmark launched to test if AI models can truly make drug discoveries (August 14, 2026)
- Insilico Medicine — DDD Benchmark portal (August 2026)
- News-Medical — Insilico Medicine nominates powerful oral drug candidate for the management of heart risks (August 12, 2026)