Answers · By Hugh Donatello, Atlas Bio · Updated 2026-10-05

How to Evaluate an AI Drug-Discovery Vendor

How should I evaluate an AI drug-discovery vendor? Ask five things: the out-of-distribution number next to the headline number, one prediction that was dated before its outcome, one documented miss, the exact decision the model is qualified for, and the base rate that decision already achieves without a model. A vendor who cannot supply all five is selling a demonstration, not a capability.

Start with the base rate you are trying to beat

Every claim of improvement needs a denominator. Wong, Siah and Lo (Biostatistics 2019) analysed 406,038 clinical trial entries covering more than 21,143 compounds from 2000 to late 2015 and found an oncology success rate of 3.4% in their sample, against 5.1% in prior widely cited studies - and considerable movement over time, with the rate recovering to 8.3% in 2015 after falling to 1.7% in 2012. They also found that trials using biomarkers for patient selection had higher overall success probabilities.

Hwang and colleagues (JAMA Internal Medicine 2016) followed 640 novel therapeutics entering pivotal trials between 1998 and 2008: 344 (54%) failed and 230 (36%) were approved. Among the failures, 195 (57%) failed for inadequate efficacy, 59 (17%) for safety, and 74 (22%) for commercial reasons.

Those three numbers frame the whole conversation. A vendor promising to fix candidate selection is addressing the 57%. Nothing in a molecular model addresses the 22% that fail commercially, and a vendor who does not say so has not thought about where their tool sits.

The questions that separate a model from a deck

Most vendor evaluations fail because the buyer asks about performance and accepts a single number. Performance is a relationship between a model and a distribution, so each question below is really asking which distribution the number came from.

What the early clinical evidence actually shows

There is now a first read on whether AI-originated molecules behave differently in the clinic. Jayatunga and colleagues (Drug Discovery Today 2024) analysed the clinical pipelines of AI-native biotechnology companies and reported an 80-90% phase 1 success rate for AI-discovered molecules, substantially above historic industry averages, with a phase 2 success rate of roughly 40% - comparable to historic averages, and on a limited sample size.

Read carefully, that is an encouraging but narrow result. Phase 1 largely tests whether a molecule is drug-like and tolerable, which is close to what these models are optimised for. Phase 2 tests whether the biological hypothesis was right, and there the advantage has not yet appeared. A vendor citing this work as proof that AI de-risks development is over-reading its own best evidence.

Treat the sample size caveat as load-bearing. These are small, young, self-selected pipelines, and attrition statistics computed on surviving programmes are the oldest trap in the field.

If the output will touch a regulatory submission

The United States Food and Drug Administration's January 2025 draft guidance on artificial intelligence to support regulatory decision-making for drug and biological products organises credibility around the context of use: the specific role and scope of the model in addressing the question of interest, assessed through a risk-based credibility assessment framework with a documented plan and report.

The operative consequence for a buyer is that credibility is never a property of a model in general. It is established for one narrow job, at one level of risk, and has to be re-established when the job changes. Ask the vendor to write down the context of use in a single sentence. If they cannot, the model is not ready to be near a submission.

Red flags

None of these are proof of a bad vendor. All of them are reasons to keep asking.

Atlas Bio publishes the artifacts this page tells you to ask for: an L1, L2 or L3 evidence tag on every claim, in-sample performance reported next to out-of-distribution performance, a multi-node reasoning audit run on every output before a human reviewer sees it, and its own documented wrong predictions kept on the pre-registration ledger.

What single question is the most diagnostic?

Ask for a prediction that was wrong, with the original record and the recalibration. It tests pre-registration, out-of-distribution honesty and institutional culture in one question, and it is the question vendors are least prepared for.

Should I ask for the training data?

Ask for its composition and provenance rather than the data itself: which assays, which species, which years, how labels were assigned, and the class balance. Composition tells you where the model will fail, and most vendors can share it without exposing proprietary work.

Is a published, peer-reviewed model better than a proprietary one?

Not automatically, but publication exposes the evaluation protocol to people with an incentive to find the flaw. If a model is proprietary, the substitute is an independent reviewer granted access under a non-disclosure agreement, and a reporting standard the work is held to.

  1. Wong CH, Siah KW, Lo AW. Estimation of clinical trial success rates and related parameters. Biostatistics 2019 (PMID 29394327)
  2. Hwang TJ et al. Failure of Investigational Drugs in Late-Stage Clinical Development and Publication of Trial Results. JAMA Internal Medicine 2016 (PMID 27723879)
  3. Jayatunga MKP et al. How successful are AI-discovered drugs in clinical trials? A first analysis and emerging lessons. Drug Discovery Today 2024 (PMID 38692505)
  4. Collins GS et al. TRIPOD+AI statement: updated guidance for reporting clinical prediction models that use regression or machine learning methods. BMJ 2024 (PMC11019967)
  5. FDA. Considerations for the Use of Artificial Intelligence To Support Regulatory Decision-Making for Drug and Biological Products. Draft guidance, January 2025

Talk to Atlas Bio.

Sponsors, researchers and clinicians: tell us the decision you are trying to make and we will show what the platform can and cannot answer.

Contact Us →