Methodology · M3

Cross-Trial Validation.

A single parameter set is tested against four independent trials — without per-trial tuning. Predictions stay within ±5% for Gr3 TRAE, ±1–2 percentage points for fatal TRAE, ±5% for ORR, and ±2 months for mPFS. The hardest test isn't matching a positive trial; it's reproducing a negative one.

One framework, four trials.

TrialIndicationPopulationN (safety)Outcome
CLEAR (Motzer 2021)RCC 1LFit, ECOG 0-1, median 62352Calibration anchor · positive
KN-775 (Makker 2022)Endometrial 2LPost-platinum, pelvic-RT in 30%406Validation · positive (fragile)
LEAP-002 (Llovet 2023)HCC 1LCirrhotic CP-A395Validation · negative · hardest test
LEAP-012 (Llovet 2024)HCC 1L + TACECirrhotic CP-A + locoregional~241Validation · positive (locoregional rescue)

Within tolerance on all four.

TrialGr3 ΔFatal ΔORR ΔmPFS Δ
CLEAR+1.3% match+1.4% (over-pred*)+2.0% match+1.4 mo match
KN-775+1.5% match (v2.0)+0.15% match−1.7% match+0.1 mo match
LEAP-002−5.1% match+2.2% (over-pred*)−0.9% match (v2.0)−0.2 mo match (v2.0)
LEAP-012−15% (no TACE v1) / match (v2.0)+1.2% match+3.0% match (v2.0)+1.1 mo match (v2.0)

* Over-predictions corrected in v2.0 via extra-fit downward coefficient and CP-A reserve split.

LEAP-002 — the hardest test.

The framework's hardest test wasn't reproducing CLEAR or KN-775 — both were positive trials. It was reproducing LEAP-002, which was negative for OS despite Atlas Bio's v1.0 predicting it would be positive (synergy coef 0.30 transferred from RCC).

The trial failed. The framework was wrong. v2.0 recalibration to s = 0.10 for HCC reproduces the negative outcome.

This is the lesson: synergy coefficients reflect tissue-specific immune-vascular biology, not portable drug-class properties. Cirrhotic immune dysfunction collapses IO contribution. The lesson is now embedded in the calculator's logic.

Reproducing a negative trial is harder than reproducing a positive one because the failure modes are more varied (any of dose, drug, population, design, indication can make a trial fail). A model that only ever "predicts positive" is useless. Atlas Bio publishes its negative predictions explicitly.

One parameter set, four trials.

The framework's calibration set is CLEAR (the first trial). The other three trials are validation cohorts — the population fragility coefficient, the substrate Gr3-shift, the CP-A reserve split, and the synergy coefficients were not re-fit per trial after the data came in.

The v2.0 refinements that close gaps (e.g., the substrate Gr3-shift for KN-775, the CP-A reserve split for LEAP-002) are mechanistic additions, not free parameters fit to the answer. Each refinement is a single-direction coefficient with a biological justification, not a flexible knob.

Future validations will use the v2.0 parameter set as-is. The next reckoning will be the next IO+TKI phase 3 trial in a new population. If the prediction holds, the framework's generalization is real. If it breaks, that's data for v3.

Want the full validation table?

27-row cross-trial table with all safety and efficacy metrics, modifiers applied, and refinement audit trail. Available to verified reviewers.

Request access →