The base rates you are predicting against
Wong, Siah and Lo (Biostatistics 2019) estimated success rates from 406,038 clinical trial entries covering more than 21,143 compounds between 2000 and late 2015. Their headline for oncology - a 3.4% overall success rate, against 5.1% in prior widely cited work - shows how much these numbers move with the sample and the method. The same analysis found the oncology rate fell to 1.7% in 2012 and recovered to 8.3% by 2015, and that trials using biomarkers for patient selection had higher overall success probabilities.
Any prediction of phase 3 success is a prediction about a shifting denominator. A model trained on a decade that included the 2012 trough and the 2015 recovery has learned something about both drug development and the funding and design practices of those years.
What published models actually achieve
The strongest public benchmark is HINT (Fu et al., Patterns 2022), which encodes drug molecule, target disease and trial eligibility criteria, adds pharmacokinetic and historical trial knowledge, and connects them in a hierarchical interaction graph. It was trained and validated on 1,160 phase 1, 4,449 phase 2 and 3,436 phase 3 trials, and achieved F1 scores of 0.665, 0.620 and 0.847 on separate test sets of 627 phase 1, 1,653 phase 2 and 1,140 phase 3 trials.
Read the ordering carefully. Phase 3 is the easiest of the three to predict, not the hardest, because by then the drug has survived two rounds of selection, the indication is narrow, the design is public and the class has precedent. The high phase 3 number is partly a statement about how much information is already on the table - which is also why it is the phase where a model adds the least that an experienced reviewer would not.
Why phase 3 trials actually fail
Hwang and colleagues (JAMA Internal Medicine 2016) tracked 640 novel therapeutics entering pivotal trials between 1998 and 2008 through 2015. 344 (54%) failed during development and 230 (36%) were approved. Of the failures, 195 (57%) were due to inadequate efficacy, 59 (17%) to safety concerns, and 74 (22%) to commercial reasons.
That distribution is the ceiling on any biology-only model. Roughly a fifth of late-stage failures are a sponsor decision about markets, financing or portfolio priority, and no feature set drawn from molecules, mechanisms and trial design contains that information. An honest phase 3 predictor states which slice of the failure distribution it addresses.
The same study found only about 40% of trial results for failed agents were published in peer-reviewed journals. Every model trained on the literature therefore learns from a sample in which successes are systematically better documented than failures.
The leakage trap, which is specific and avoidable
Trial-outcome prediction is unusually exposed to temporal contamination because the registry is continuously updated. Enrolment status, results postings, label changes and the publication itself all appear in public records after the outcome is known, and any of them can creep into a feature set as a near-perfect predictor of the answer.
Kapoor and Narayanan (Patterns 2023) found leakage across 17 fields and 294 papers, with a taxonomy of eight distinct types - and in their own reproduction, correcting for it collapsed the apparent advantage of complex models over logistic regression. A phase 3 predictor should therefore be evaluated on a strict temporal split, trained only on features that existed before the prediction date, and reported with the feature-availability date stated.
The practical test for a buyer is simple: ask what the model knew on the day it made the prediction, and ask for a prediction made before a readout that has since occurred.
How to use a probability that is honest
Calibration matters more than discrimination for this use. A model with a modest AUC whose 70% predictions come true about 70% of the time supports portfolio arithmetic; a sharper model that is systematically overconfident does not, because every decision it informs is sized wrongly.
Use the output to rank where to spend diligence, to identify which phase 2 signals need strengthening before the phase 3 design locks, and to force an explicit conversation about the top risk-driving features. Do not use it as a go or no-go verdict on a programme, and do not accept a point estimate without an interval.
- Require a temporal split, not a random one
- Require calibration curves alongside AUC or F1
- Require the prediction to be dated and locked before the readout
- Require the model's out-of-scope list: manufacturing, competitive landscape, sponsor execution, unanticipated rare safety events
- Treat any model trained on published literature as optimistic, because failures are under-published
Atlas Bio's Phase 3 Outcome Predictor forecasts gene-therapy phase 3 outcomes from publicly available phase 1 and 2 data across eight feature categories, is validated on a temporal split with no future-data leakage, outputs a probability with explicit confidence intervals rather than a binary verdict, and publishes what it does not model: manufacturing failures, competitive landscape changes, sponsor-specific execution risk and black-swan safety events.
Does a high AUC mean a phase 3 predictor is trustworthy?
No. AUC measures ranking, not whether the stated probabilities are right, and it is the quantity most inflated by leakage. Ask for calibration on a temporally held-out set and for predictions issued before the outcomes they claim.
Can a model predict a trial that fails for commercial reasons?
Not from biology. Published analysis attributes about 22% of late-stage failures to commercial decisions. A model can only reach those if sponsor-level and market features are explicitly in the feature set, which carries its own leakage and fairness problems.
Is phase 2 or phase 3 harder to predict?
Phase 2. In the HINT benchmark, phase 2 scored lowest of the three phases. Phase 2 is where the biological hypothesis is genuinely tested, so there is less prior evidence to lean on and more of the outcome is irreducibly uncertain.
- Wong CH, Siah KW, Lo AW. Estimation of clinical trial success rates and related parameters. Biostatistics 2019 (PMID 29394327)
- Fu T, Huang K, Xiao C, Glass LM, Sun J. HINT: Hierarchical interaction network for clinical-trial-outcome predictions. Patterns 2022 (PMC9024011)
- Hwang TJ et al. Failure of Investigational Drugs in Late-Stage Clinical Development and Publication of Trial Results. JAMA Internal Medicine 2016 (PMID 27723879)
- Kapoor S, Narayanan A. Leakage and the reproducibility crisis in machine-learning-based science. Patterns 2023 (PMC10499856)
Talk to Atlas Bio.
Sponsors, researchers and clinicians: tell us the decision you are trying to make and we will show what the platform can and cannot answer.
Contact Us →