Prediction and postdiction are different claims
The distinction is not pedantic. Nosek and colleagues (PNAS 2018) define it precisely: postdiction is characterised by the use of data to generate hypotheses about why something occurred, and prediction by the acquisition of data to test ideas about what will occur. Both are legitimate science. Only one tells you whether a model works.
The same authors state the failure mode plainly: presenting postdictions as predictions can increase the attractiveness and publishability of findings by falsely reducing uncertainty, and ultimately decreases reproducibility. Their proposed remedy is procedural rather than moral - preregistration of an analysis plan is committing to analytic steps without advance knowledge of the research outcomes.
In computational biology this is the difference between a vendor saying 'our model explains why that trial failed' and 'our model said so in writing, eleven months before the readout'. Only the second is a track record.
The evidence that locking a prediction changes the answer
Pre-registration is one of the rare methodological reforms with a measured effect size. Kaplan and Irvin (PLoS ONE 2015) examined large National Heart, Lung, and Blood Institute trials of drugs and dietary supplements. Among studies published before 2000, 17 of 30 (57%) reported a significant benefit on the primary outcome. Among the 25 trials published after 2000, when prospective registration of the primary outcome at ClinicalTrials.gov became the norm, only 2 (8%) did.
The authors also noted that 12 of the 25 registered trials reported significant positive effects on cardiovascular variables other than the pre-specified primary outcome - which is what the pre-2000 literature would have headlined. Nothing about the biology changed in 2000. What changed was the ability to choose the outcome after seeing the data.
Registration is also not self-enforcing. Anderson and colleagues (NEJM 2015) examined 13,327 trials subject to United States results-reporting requirements and completed between 2008 and 2012, and found that only 13.4% reported summary results within twelve months of completion, and 38.3% at any time up to the September 2013 cut-off. A registry entry nobody checks is paperwork; a registry entry someone audits is evidence.
Why machine learning needs this more than wet biology does
A wet-lab result is expensive to re-run, so it tends to be run once, honestly. A model can be re-run a thousand times in an afternoon, each run a chance to find a split, a threshold or a feature set that flatters the result. Nothing in that loop is dishonest, and the end product is still unreliable.
Kapoor and Narayanan (Patterns 2023) surveyed machine-learning-based science and found leakage - training information contaminating evaluation - reported in 17 fields, collectively affecting 294 papers, in some cases producing wildly overoptimistic conclusions. They catalogue eight distinct types, from textbook train-test contamination to open research problems. In their own reproduction of a civil-war prediction literature, correcting the errors left complex models performing no better than decades-old logistic regression.
A pre-registered prediction is the cheapest available defence, because it fixes the evaluation before the modeller can see which evaluation would be flattering.
- The exact quantity being predicted, with units and a measurement source
- The point estimate and an interval, not a direction
- The analysis that will decide hit or miss, written before the data arrive
- The kill criterion: what result would make the team abandon the hypothesis
- A timestamp and a hash, so the record cannot be quietly edited
- The validation event and its expected date
What pre-registration does not buy you
It does not make predictions correct. A locked prediction can be badly wrong, and the discipline is worth nothing unless the wrong ones stay visible. The strongest evidence that a group pre-registers is a preserved miss with the recalibration documented next to it.
It does not prevent selective reveal either. A team that locks fifty predictions and publishes the six that landed has gamed the system exactly as thoroughly as one that never locked anything. The ledger has to be enumerable: every lock, its outcome, and the ones still open.
And it does not substitute for out-of-distribution testing. A model can be pre-registered and still be evaluated only on candidates that look like its training set. Lock the prediction and report the generalisation gap; neither alone is sufficient.
Atlas Bio pre-registers its predictions: hypotheses, primary endpoints, the statistical analysis plan and kill criteria are SHA-256 hashed and timestamped before data collection, the hash is committed to a ledger, and wrong predictions are preserved on that ledger alongside the correct ones.
Is pre-registration the same as publishing a protocol?
A protocol describes what you will do. A pre-registration additionally fixes what you predict will happen and what would count as a failure, before the data exist. The prediction and the kill criterion are the parts that make a later claim of accuracy checkable.
How can a computational prediction be timestamped credibly?
Hash the locked document and commit the hash and time to a record that cannot be edited afterwards. On reveal, the original file is published and the hash is recomputed and compared. Anyone can verify that the prediction was not adjusted after the result arrived.
Does pre-registration slow research down?
It slows down claims, not work. Exploratory analysis stays unrestricted and is reported as exploratory. What is blocked is the free upgrade of an exploratory finding into a confirmed prediction once the outcome is already known.
- Nosek BA, Ebersole CR, DeHaven AC, Mellor DT. The preregistration revolution. PNAS 2018 (PMC5856500)
- Kaplan RM, Irvin VL. Likelihood of Null Effects of Large NHLBI Clinical Trials Has Increased over Time. PLoS ONE 2015
- Anderson ML et al. Compliance with results reporting at ClinicalTrials.gov. New England Journal of Medicine 2015 (PMC4508873)
- Kapoor S, Narayanan A. Leakage and the reproducibility crisis in machine-learning-based science. Patterns 2023 (PMC10499856)
Talk to Atlas Bio.
Sponsors, researchers and clinicians: tell us the decision you are trying to make and we will show what the platform can and cannot answer.
Contact Us →