One framework, four trials.
| Trial | Indication | Population | N (safety) | Outcome |
|---|---|---|---|---|
| CLEAR (Motzer 2021) | RCC 1L | Fit, ECOG 0-1, median 62 | 352 | Calibration anchor · positive |
| KN-775 (Makker 2022) | Endometrial 2L | Post-platinum, pelvic-RT in 30% | 406 | Validation · positive (fragile) |
| LEAP-002 (Llovet 2023) | HCC 1L | Cirrhotic CP-A | 395 | Validation · negative · hardest test |
| LEAP-012 (Llovet 2024) | HCC 1L + TACE | Cirrhotic CP-A + locoregional | ~241 | Validation · positive (locoregional rescue) |
Within tolerance on all four.
| Trial | Gr3 Δ | Fatal Δ | ORR Δ | mPFS Δ |
|---|---|---|---|---|
| CLEAR | +1.3% match | +1.4% (over-pred*) | +2.0% match | +1.4 mo match |
| KN-775 | +1.5% match (v2.0) | +0.15% match | −1.7% match | +0.1 mo match |
| LEAP-002 | −5.1% match | +2.2% (over-pred*) | −0.9% match (v2.0) | −0.2 mo match (v2.0) |
| LEAP-012 | −15% (no TACE v1) / match (v2.0) | +1.2% match | +3.0% match (v2.0) | +1.1 mo match (v2.0) |
* Over-predictions corrected in v2.0 via extra-fit downward coefficient and CP-A reserve split.
LEAP-002 — the hardest test.
The framework's hardest test wasn't reproducing CLEAR or KN-775 — both were positive trials. It was reproducing LEAP-002, which was negative for OS despite Atlas Bio's v1.0 predicting it would be positive (synergy coef 0.30 transferred from RCC).
The trial failed. The framework was wrong. v2.0 recalibration to s = 0.10 for HCC reproduces the negative outcome.
Reproducing a negative trial is harder than reproducing a positive one because the failure modes are more varied (any of dose, drug, population, design, indication can make a trial fail). A model that only ever "predicts positive" is useless. Atlas Bio publishes its negative predictions explicitly.
One parameter set, four trials.
The framework's calibration set is CLEAR (the first trial). The other three trials are validation cohorts — the population fragility coefficient, the substrate Gr3-shift, the CP-A reserve split, and the synergy coefficients were not re-fit per trial after the data came in.
The v2.0 refinements that close gaps (e.g., the substrate Gr3-shift for KN-775, the CP-A reserve split for LEAP-002) are mechanistic additions, not free parameters fit to the answer. Each refinement is a single-direction coefficient with a biological justification, not a flexible knob.
Future validations will use the v2.0 parameter set as-is. The next reckoning will be the next IO+TKI phase 3 trial in a new population. If the prediction holds, the framework's generalization is real. If it breaks, that's data for v3.
Want the full validation table?
27-row cross-trial table with all safety and efficacy metrics, modifiers applied, and refinement audit trail. Available to verified reviewers.
Request access →