Biotech · 2 min read

Beyond Memorization: Insilico's DDD Benchmark Tests Whether AI Can Actually Discover Drugs

Insilico Medicine launches DDD Benchmark as a Service — decontaminated real-world drug-discovery evaluation for frontier AI models, with 300+ foundation tests and proprietary candidate baselines.

By Classy AI News · July 31, 2026

Beyond Memorization: Insilico's DDD Benchmark Tests Whether AI Can Actually Discover Drugs

Drug-discovery AI benchmarks have a contamination problem: models memorize test questions from training data and score well on paper, then fail when a chemist asks for a real preclinical candidate decision. On July 30, 2026, Insilico Medicine (HKEX:3696) launched Drug Discovery and Development (DDD) Benchmark as a Service — a standardized evaluation framework built on decontaminated real-world data and proprietary validated programs.

Scientific workspace with research materials

Can these models actually discover drugs?

Insilico's framing is blunt: "Can these models actually discover drugs, or are they just good at taking tests?" Most public benchmarks, the company argues, suffer from severe data contamination — inflated scores that collapse under real candidate-selection pressure.

The DDD Benchmark anchors evaluation in carefully decontaminated real-world data and more than 30 validated preclinical candidate programs drawn from Insilico's own pipeline — including Rentosertib (ISM001-055), a first-in-class TNIK inhibitor now in Phase III for IPF.

Two complementary suites

Drug Discovery Foundations — more than 300 evaluations from proprietary out-of-distribution test sets and rigorously decontaminated public data. Core competencies span disease biology, molecular property prediction, retrosynthesis, structure-based design, and clinical development.

Drug Candidate Essentials — end-to-end decision-making tests using proprietary reference baselines from Insilico's validated programs.

Organizations submit models via a standard chat-completions API; Insilico scores outputs against expert baselines and delivers verified scorecards. Public leaderboard placement is optional.

Research laboratory environment with equipment

Why BaaS matters now

As frontier general-purpose models infiltrate pharma R&D, buyers need independent verification — not leaderboard vanity metrics. Insilico reports compressing preclinical-candidate nomination timelines to roughly 12–18 months versus 2.5–4+ years traditionally, with 10+ IND clearances across its AI-designed portfolio.

The service is live at dddbench.insilico.com, alongside Insilico's broader benchmark explorer at ddb.insilico.com covering 217 benchmarks across biology, chemistry, materials, and longevity.

For pharma AI, July 30 marks a shift from "trust our demo" to pay-for-proof — with contamination controls that public academic benchmarks rarely enforce.

### Sources

Newsletter

Get the dispatch

One field. One email when we publish. Privacy.