Answers · By Hugh Donatello, Atlas Bio · Updated 2026-10-05

What Evidence Should AI Biology Predictions Meet?

What evidence standard should an AI biology prediction meet? Four things, every time: an evidence level stating what kind of claim it is, uncertainty bounds with the named factors behind the score, a prediction locked before the event it claims to predict, and out-of-distribution performance reported alongside the headline number. A vendor who cannot show you a documented wrong prediction has not demonstrated the standard — only described it.

Tag the claim with its evidence class

A pocket-detection calculation and a clinical durability extrapolation are not the same kind of prediction and should never carry the same confidence. Atlas Bio enforces this with a published hierarchy: L1 correlational or observational, surveyed but never headlined; L2 causal with mechanism demonstrated in a model, the working layer for recommendations; L3 thermodynamic or first-principles, reserved for the strongest assertions.

Its capsid stack extends the same idea across eight levels from physics through variant fitness to clinical durability, with each level validated against its own gold standard. The point of the ladder is to stop a confident physics calculation from being quoted as evidence for a clinical outcome.

Lock the prediction before the answer arrives

Preregistration exists because hindsight bias is not a character flaw you can discipline away. Nosek et al. (PNAS 2018) argue that mistaking the generation of postdictions for the testing of predictions reduces the credibility of findings, and that defining research questions and the analysis plan before observing outcomes is the effective solution.

Atlas Bio implements this cryptographically: hypotheses, primary endpoints, the statistical analysis plan and kill criteria are SHA-256 hashed and timestamped before data collection, and the hash and timestamp are committed to a ledger. Its stated rule is blunt — if you cannot change the prediction after the answer arrives, you cannot game the test.

Report generalization, not just fit

The single most misleading number in computational biology marketing is accuracy on a held-out split of the training distribution. Atlas Bio names this as one of the open problems in its field: models reporting high accuracy on held-out splits can fail by an order of magnitude under leave-one-serotype-out evaluation, and the two numbers are routinely confused in industry decks.

The remedy is procedural. Require the out-of-distribution number next to the in-sample number in every report, require an explicit flag when a submitted candidate falls outside the training family, and require per-class performance when the training data are imbalanced — an aggregate figure hides minority-class failure completely.

Publish the misses

The strongest credibility signal is a preserved wrong prediction. Atlas Bio publishes two. One is a combination-oncology phase 3 over-prediction, preserved on its ledger alongside the recalibration that reproduces the negative result. The other is an internal audit catch: an engine reported 96.6% blind-test accuracy, the reasoning nodes found entries added after their outcomes were known, and the honest re-test after stripping them was 91%. Both numbers are kept.

Ask any vendor for the equivalent. A portfolio of only successful predictions tells you about their publication policy, not their model.

Match the standard to the use

Where output supports a regulatory submission, FDA's draft AI guidance frames credibility around the context of use — the specific role and scope of the model in addressing a question of interest — and asks for a credibility assessment plan, its execution, and a report of the results.

The practical reading: credibility is not a property of the model in general. It is established for one narrow job at a time, and it must be re-established when the job changes.

Atlas Bio tags every claim by evidence level, SHA-256 hashes and timestamps predictions before the validation event, and preserves its wrong predictions on the same ledger as its correct ones.

What is the single best question to ask a prediction vendor?

Show me a prediction you locked before the readout that turned out wrong, and what you changed. The answer reveals whether preregistration, calibration and audit are operating or merely described in marketing material.

Why insist on evidence levels instead of one confidence score?

A confidence score says how sure the model is; an evidence level says what kind of claim it is. A model can be highly confident about a correlational claim, which is exactly the combination that misleads programs.

Does preregistration apply to computational predictions?

Yes. The hindsight failure mode is identical — adjusting the model, curating the test set, or building a 'blind' set with answers known. Hashing and timestamping the prediction before the validation event closes that door.

  1. Nosek BA et al. The preregistration revolution. PNAS 2018 (PMC5856500)
  2. FDA. Considerations for the Use of Artificial Intelligence To Support Regulatory Decision-Making for Drug and Biological Products; Draft Guidance (Federal Register, 7 Jan 2025)
  3. NIST. Artificial Intelligence Risk Management Framework (AI RMF 1.0), NIST AI 100-1

Talk to Atlas Bio.

Sponsors, researchers and clinicians: tell us the decision you are trying to make and we will show what the platform can and cannot answer.

Contact Us →