Methodology & Prediction Framework
Prediction baseline is mechanistic / class-based. Combo prediction departs from the additive logic used for safety — for efficacy we use max(p_IO, p_TKI) × (1 + s) where the synergy coefficient s reflects mechanism complementarity (immune priming via VEGFR blockade, vascular normalization enabling T-cell infiltration). For mPFS and mOS, the combo is anchored to the better-arm baseline (not the sum) and amplified by a tumor-specific synergy multiplier. Population fitness modifier applied for ECOG, age, prior therapy lines, organ reserve.
Pembrolizumab class baseline (mono, 200 mg Q3W)
Pan-tumor unselected ORR 15–25% · Biomarker-enriched (PD-L1≥50%, MSI-H, TMB-high) ORR 35–55%
Median PFS 3–6 mo unselected · 8–12 mo enriched · DCR 40–55%
Median OS 12–18 mo unselected · 24–32 mo enriched · 5-yr OS 16–23%
Time-to-response median 8–12 wk · Duration of response 12–24 mo
Lenvatinib class baseline (24 mg; 18 / 12-8 mg scaling)
Tumor-specific ORR: DTC 65% · HCC 24% · endometrial 14% · thymic 38%
Median PFS: DTC 18.3 mo · HCC 7.4 mo · endometrial 5.6 mo
Median OS: DTC NR (delayed crossover) · HCC 13.6 mo · endometrial 10.6 mo
Time-to-response 6–10 wk · Mostly partial responses, low CR rate
Combination Efficacy + Risk/Benefit + RWE Calculator v2.3 ✓ validated × 4 trials
Input setting + indication + dose + population fitness + biomarker → outputs predicted combo ORR (with CI), median PFS, median OS, depth of response score, durability index, time-to-response, risk/benefit pairing vs control (PFS gained, NNT, benefit density, automated verdict). v2.3 adds the Clinical Trial / Real-World Practice toggle applying literature-derived efficacy attenuation factors (ORR/PFS/OS ×0.85) — on top of v2.1 risk/benefit module and v2.0 refinements (Child-Pugh-stratified subsequent therapy modeling, biomarker enrichment, line-of-therapy modifier). Use preset buttons to reproduce validation trials.
Inputs
Predicted combination efficacy
Predicted depth, durability, kinetics v2.1
Risk / Benefit pairing (vs control)
Predicted response-determining contributors
| Component | Driver | ORR contribution | PFS contribution |
|---|
Population Response Grid — Same Regimen, Different Patients v2.5
100 virtual patients per population, sorted by depth of response. The visual mass of green squares is the responder-rate story — no chart-reading required. Default comparison: fit RCC (CLEAR) vs fragile post-platinum endometrial (KN-775), showing the framework's central efficacy thesis: biomarker-aligned indication dominates dose, drug interaction, and even arm choice for absolute response rate.
in Profile A
Same dose.
Multi-Combination Efficacy Comparator v2.2
Same patient profile evaluated across all approved IO+TKI / IO+antiangiogenic combinations for the selected indication. Pulls inputs from the calculator above. Patient fitness modifier (kfit) applied to ORR and PFS uniformly across combos. Recommended combo (best benefit density: PFS gain per unit fatal risk, drawing from the safety analysis) marked ★. Efficacy-first ranking available via the toggle below.
Decision logic
Best combo for efficacy = highest ORR × DCR × (mPFS / 12) — favors deep, durable, fast responses.
Tie-breaker priority: (1) longest mPFS, (2) longest mOS, (3) highest CR rate, (4) regulatory status, (5) familiarity / payer access.
Combos covered
RCC 1L: Pembro+Lenva (CLEAR), Nivo+Cabo (CheckMate-9ER), Pembro+Axi (KN-426)
Endometrial 2L: Pembro+Lenva (KN-775) — only approved IO+TKI; vs Pembro+chemo (NRG-GY018) in MMRp 1L
HCC 1L: Atezo+Bev (IMbrave150) only positive; Pembro+Lenva (LEAP-002) negative; Durva+Treme (HIMALAYA) approved
Part 1 — Pembrolizumab Monotherapy: Paper vs Class Prediction
Eight pivotal monotherapy trials evaluated against the mechanistic anti-PD-1 efficacy baseline. The class baseline distinguishes biomarker-enriched populations (PD-L1≥50%, MSI-H, TMB-high, melanoma class) from unselected populations.
1.1 KEYNOTE-006 Melanoma 1L · N=556 (Q3W arm) · vs ipilimumab
| Metric | Predicted (melanoma class) | Observed | Δ | Interpretation |
|---|---|---|---|---|
| ORR | 30–38% | 33.7% | match | melanoma is reference IO-responsive class |
| CR rate | 5–8% | 6.7% | match | — |
| Median PFS | 4.5–6 mo | 5.5 mo | match | — |
| Median OS | 28–34 mo | 32.7 mo | match | vs ipi 15.9 mo (HR 0.68) |
| 3-yr OS | 40–48% | 50% | slight ↑ | tail-of-curve advantage |
| 5-yr OS | 30–38% | 38.7% | match | plateau reached |
| Median DoR | not reached (NR) | NR (≥3 yr) | match | durable plateau |
| Time to response | 10–14 wk | 12 wk | match | — |
1.2 KEYNOTE-024 NSCLC 1L PD-L1≥50% · N=154 · vs platinum chemo
| Metric | Predicted (PD-L1 enriched) | Observed | Δ | Interpretation |
|---|---|---|---|---|
| ORR | 40–50% | 44.8% | match | biomarker-enriched performs as predicted |
| CR rate | 2–5% | 2.6% | match | NSCLC has lower CR rate than melanoma |
| Median PFS | 9–12 mo | 10.3 mo | match | vs chemo 6.0 mo (HR 0.50) |
| Median OS | 22–28 mo | 30.0 mo | +2–8 mo over | crossover-protected outperformance |
| 5-yr OS | 25–32% | 31.9% | match | — |
| Median DoR | NR (long tail) | NR (median follow-up 5+ yr) | match | — |
1.3 KEYNOTE-042 NSCLC 1L PD-L1≥1% · N=636 · vs platinum chemo
| Metric | Predicted (broader PD-L1) | Observed | Δ | Interpretation |
|---|---|---|---|---|
| ORR | 22–30% | 27.3% | match | broader pool dilutes biomarker signal |
| Median PFS | 5–7 mo | 5.4 mo | match | vs chemo 6.5 (PFS not separated in TPS 1–49% subset) |
| Median OS | 15–18 mo | 16.7 mo | match | vs chemo 12.1 (HR 0.81 overall) |
| OS in TPS≥50% subset | 20–28 mo | 20.0 mo | match | recapitulates KN-024 |
| OS in TPS 1–49% subset | 13–17 mo | 13.4 mo | match | marginal benefit, expected |
1.4 KEYNOTE-045 Urothelial 2L · N=266 · vs taxane/vinflunine
| Metric | Predicted (unselected urothelial) | Observed | Δ | Interpretation |
|---|---|---|---|---|
| ORR | 18–25% | 21.1% | match | moderate IO responsiveness |
| CR rate | 5–9% | 9.3% | slight ↑ | urothelial CRs unusual depth |
| Median PFS | 2–3 mo | 2.1 mo | match | not vs chemo PFS (3.3 mo) |
| Median OS | 9–12 mo | 10.3 mo | match | vs chemo 7.4 (HR 0.73) — survival benefit despite weak PFS |
| Median DoR | NR (long tail expected) | NR (median follow-up 4 yr) | match | — |
1.5 KEYNOTE-204 Classical Hodgkin r/r · N=148 · vs brentuximab vedotin
| Metric | Predicted (Hodgkin class — IO-hypersensitive) | Observed | Δ | Interpretation |
|---|---|---|---|---|
| ORR | 60–72% | 65.6% | match | Hodgkin is most IO-responsive solid tumor analog |
| CR rate | 20–30% | 24.5% | match | — |
| Median PFS | 11–15 mo | 13.2 mo | match | vs brentuximab 8.3 mo (HR 0.65) |
| Median DoR | 14–22 mo | 20.5 mo | match | — |
| Median OS | NR | NR | match | — |
| 2-yr OS | 85–92% | 89.4% | match | — |
1.6 KEYNOTE-158 MSI-H tissue-agnostic basket · N=351
| Metric | Predicted (MSI-H biomarker-enriched) | Observed | Δ | Interpretation |
|---|---|---|---|---|
| ORR | 30–42% | 34.3% | match | MSI-H = TMB surrogate |
| CR rate | 8–14% | 10.0% | match | — |
| Median PFS | 3.5–6 mo | 4.1 mo | match | median misleading; long tail |
| Median OS | 20–28 mo | 23.5 mo | match | — |
| Median DoR | NR | NR (median follow-up 13 mo) | match | — |
| 4-yr OS (responders) | ~70% | ~73% | match | — |
1.7 KEYNOTE-240 HCC 2L · N=278 · vs placebo
| Metric | Predicted (HCC class — modest IO) | Observed | Δ | Interpretation |
|---|---|---|---|---|
| ORR | 15–22% | 18.4% | match | cirrhotic immune-context blunting |
| Median PFS | 2.5–3.5 mo | 3.0 mo | match | vs placebo 2.8 (HR 0.78) |
| Median OS | 13–17 mo | 13.9 mo | match | vs placebo 10.6 (HR 0.78); P=0.0238 missed pre-spec α |
| Median DoR | 10–18 mo | 13.8 mo | match | — |
| Statistical positivity | borderline | missed (P>0.0174) | trial-design issue | positive HR but didn't cross α-spending boundary |
1.8 KEYNOTE-052 Urothelial 1L cisplatin-ineligible · N=370
| Metric | Predicted (unselected urothelial 1L) | Observed | Δ | Interpretation |
|---|---|---|---|---|
| ORR | 22–32% | 28.6% | match | 1L slightly better than 2L |
| ORR in CPS≥10 | 40–48% | 47.3% | match | biomarker-enriched as expected |
| CR rate | 5–10% | 8.9% | match | — |
| Median PFS | 2–3 mo | 2.3 mo | match | same shallow IO PFS pattern |
| Median OS | 10–14 mo | 11.3 mo | match | — |
| Median DoR | NR (long tail) | NR (median follow-up 56 mo) | match | 30%+ ongoing at 4 yr |
Pembro mono — efficacy prediction performance summary
| Setting | ORR performance | OS performance | Biomarker dependence |
|---|---|---|---|
| Class-predicted (melanoma, MSI-H, Hodgkin) | within ±5% | within ±10% mo | biomarker-aligned |
| NSCLC PD-L1 enriched (TPS≥50) | match | slight ↑ | strong PD-L1 dose-response |
| Urothelial (broad) | match | match | moderate; CPS-graded |
| HCC 2L | match | marginal benefit | cirrhotic blunting |
| PFS as endpoint (any indication) | unreliable | median PFS hides long DoR tail; use ORR + OS together | |
Part 2 — Lenvatinib Monotherapy: Paper vs Class Prediction
Five pivotal mono trials, dose-stratified 24 / 18 / 12-8 mg. Lenvatinib's efficacy class baseline is tumor-specific — unlike pembro's biomarker-graded behavior, lenva's response rate varies 5× across approved indications (DTC 65% vs endometrial 14%) driven by tumor-vasculature dependence and FGFR/RET co-targets.
2.1 SELECT RAI-refr DTC · 24 mg · N=261 · vs placebo
| Metric | Predicted (DTC class) | Observed | Δ | Interpretation |
|---|---|---|---|---|
| ORR | 55–70% | 64.8% | match | DTC = reference VEGFR-responsive class |
| CR rate | 1–3% | 1.5% | match | TKI characteristic — low CR, high PR |
| Median PFS | 15–22 mo | 18.3 mo | match | vs placebo 3.6 mo (HR 0.21) — most dramatic ratio in lenva data |
| Median OS | NR (delayed crossover) | NR | match | placebo crossover dilutes OS signal |
| Median DoR | not reported separately | — | — | — |
| Time to response | 4–8 wk | 2.0 mo (median) | match | fast TKI-class kinetics |
2.2 REFLECT HCC 1L · 12/8 mg · N=476 · vs sorafenib
| Metric | Predicted (HCC, 75–80% dose) | Observed | Δ | Interpretation |
|---|---|---|---|---|
| ORR (mRECIST) | 20–28% | 24.1% | match | vs sora 9.2% — dose-scaled prediction validates |
| ORR (RECIST 1.1) | 16–22% | 18.8% | match | — |
| CR rate | 1–3% | 1.5% | match | — |
| Median PFS | 6–9 mo | 7.4 mo | match | vs sora 3.7 (HR 0.66) |
| Median OS | 12–14 mo (non-inferior) | 13.6 mo | match | vs sora 12.3 (HR 0.92, NI met) — primary endpoint passed via NI margin |
| Median DoR | not separately reported | — | — | — |
2.3 Study 111 Endometrial mono · 24 mg · N=124
| Metric | Predicted (endometrial mono) | Observed | Δ | Interpretation |
|---|---|---|---|---|
| ORR | 10–18% | 14.3% | match | endometrial less TKI-responsive than DTC |
| CR rate | 1–3% | 1.6% | match | — |
| Median PFS | 4–7 mo | 5.6 mo | match | — |
| Median OS | 9–12 mo | 10.6 mo | match | — |
| Median DoR | 5–9 mo | 5.4 mo | match | — |
2.4 Study 211 DTC · 18 vs 24 mg · N=152
| Metric | Pred 18 | Obs 18 | Pred 24 | Obs 24 | Verdict |
|---|---|---|---|---|---|
| ORR | 45–55% | 40.3% | 55–70% | 57.3% | 18 mg slight ↓ |
| Median PFS | 13–17 mo | 15.2 mo | 15–22 mo | NR (truncated FU) | trend match |
| CR rate | 0–3% | 1% | 1–3% | 2% | match |
| Median DoR | not reported | — | — | — | — |
2.5 REMORA Thymic carcinoma · 24 mg · N=42
| Metric | Predicted (rare tumor — wide CI) | Observed | Δ | Interpretation |
|---|---|---|---|---|
| ORR | 25–45% (wide CI) | 38.1% | match | small N → wide CI |
| Median PFS | 6–11 mo | 9.3 mo | match | — |
| Median OS | 20–28 mo | NR (immature) | trending high | — |
| Median DoR | not reported | 7.8 mo | shorter than expected | thymic biology |
Lenva mono — efficacy prediction performance summary
| Setting | ORR performance | PFS performance | Tumor-class fit |
|---|---|---|---|
| 24 mg DTC (RAI-refr) | within ±5% | within ±2 mo | reference (highest activity) |
| 12/8 mg HCC | dose-scaling validates | within ±2 mo | vasculature-driven, cirrhotic-OK |
| 24 mg endometrial | within ±5% | match | moderate (combo-amplifiable) |
| 18 mg DTC | slight ↓ vs predicted | match | dose-step ORR loss is non-linear |
| Small-N rare tumors (REMORA) | wide CIs | match | statistical noise dominant |
Part 3 — KEYNOTE-581 / CLEAR: Paper vs Max-of-Arms × Synergy Prediction
Lenva 20 mg + pembro 200 mg Q3W, 1L advanced RCC, N=355 (combo arm). The combo prediction is anchored to the better-arm baseline (lenva mono in this case for ORR and PFS), then amplified by an RCC-specific synergy coefficient calibrated to capture VEGFR-immune crosstalk.
3.0 Combo prediction inputs
Pembro mono in RCC (KEYNOTE-427)
ORR 36.4% (cohort A, ccRCC) · mPFS 7.1 mo · mOS NR
2-yr OS 71%
Lenva mono 20 mg in RCC (interpolated)
ORR 27% (per Hutson 2010 / Phase II) · mPFS 9.0 mo · mOS 18.4 mo
Single-agent VEGFR-TKI activity in ccRCC
3.1 Headline — Predicted vs Observed
| Metric | Predicted | Observed (Motzer 2021/2024) | Δ | Note |
|---|---|---|---|---|
| ORR | 65–75% | 71.0% | match | synergy coef 0.45 calibrated correctly |
| CR rate | 10–18% | 16.1% | match | combination amplifies CR rate beyond either mono |
| Median PFS | 20–25 mo | 23.9 mo | match | vs sunitinib 9.2 (HR 0.39) — 2.6× ratio |
| Median OS | ~50 mo | 53.7 mo | match | vs sunitinib 54.3 — no OS advantage despite huge PFS |
| OS HR | 0.85–1.00 | 0.79 (NS) | match | — |
| Median DoR | 20–28 mo | 26.7 mo | match | — |
| Time to response | 1.5–2.5 mo | 1.9 mo | match | fast TKI-anchored kinetics |
3.2 Class A — ORR synergy components
| Mechanism | IO contribution | TKI contribution | Synergy bonus | Observed contribution to ORR |
|---|---|---|---|---|
| VEGFR blockade → vascular normalization → T-cell infiltration | — | direct | +8–12% | quantifiable in CD8 IHC of paired biopsies |
| Reduced MDSC / Treg infiltrate (VEGF-mediated suppression lifted) | indirect | direct | +5–8% | — |
| FGFR inhibition (lenva off-target, hypoxia-relieving) | — | direct | +3–5% | partial; FGFR-amplified subset benefits more |
| RET / KIT / PDGFR (lenva multi-kinase, broader vasculature) | — | direct | +2–3% | — |
| PD-1 axis blockade (independent immune activation) | direct | — | — | ±15% baseline |
| Total synergy bonus (modeled) | +18–28% | observed +35% over best mono — fits upper bound | ||
3.3 Class B — PFS amplification components
| Mechanism | Predicted PFS contribution | Observed | Δ | Note |
|---|---|---|---|---|
| Lenva mono baseline PFS (RCC, dose-scaled to 20 mg) | 9–11 mo | — | — | floor |
| Pembro mono baseline PFS (RCC ccRCC) | 6–8 mo | — | — | — |
| Synergy multiplier on better arm (×2.4) | 20–25 mo | 23.9 mo | match | RCC synergy coef calibrated |
| vs sunitinib control (HR 0.39) | HR 0.40–0.50 | HR 0.39 | match | — |
| 2-yr PFS | 40–48% | 49.5% | match | — |
3.4 Class C — OS dilution by subsequent therapy
| Factor | Predicted impact on OS HR | Observed in CLEAR | Note |
|---|---|---|---|
| Sunitinib arm crossover/subsequent IO | HR 0.10–0.20 dilution | ~83% sunitinib pts received subsequent IO | massive dilution of OS signal |
| Combo arm subsequent therapy effectiveness | partially limited (already received IO) | ~57% received subsequent therapy | — |
| Predicted OS HR (with crossover) | 0.85–1.00 | 0.79 (NS) | match |
| What OS would have been without crossover | HR ~0.55–0.65 | not observable | cf. KEYNOTE-426 OS HR 0.74 (less crossover) |
3.5 Class D — Depth of response & durability
| Metric | Predicted | Observed | Δ | Note |
|---|---|---|---|---|
| CR rate | 10–18% | 16.1% | match | 3× pembro mono CR (5.2%), 8× lenva mono CR (1–2%) |
| ≥30% tumor shrinkage rate | 75–85% | 83.7% | match | depth-of-response amplified |
| ≥50% tumor shrinkage rate | 50–60% | ~58% | match | — |
| Median DoR (responders) | 20–28 mo | 26.7 mo | match | — |
| Ongoing response at 24 mo | 50–60% | 58% | match | — |
| Disease control rate (DCR) | 85–92% | 89.0% | match | — |
3.6 Response forensics — what drove the 71% ORR?
| Subset | Mechanism | Estimated ORR contribution |
|---|---|---|
| IO-naive responders (would respond to pembro alone) | endogenous PD-L1 / TMB / inflammatory signature | ~25% |
| TKI-only responders (would respond to lenva alone, IO non-contributor) | VEGFR-driven vascular dependence | ~20% |
| Synergy-only responders (need both for response) | VEGFR blockade enables T-cell infiltration in cold tumors | ~22% |
| Non-responders (~29%) | resistance / poor T-cell repertoire / bypass pathways | — |
3.7 Population effect — CLEAR vs other combos
| Combo trial | Population | ORR | Synergy coef inferred |
|---|---|---|---|
| CLEAR (RCC fit) | ECOG 0–1, ccRCC, IO-naive | 71% | 0.45 (calibration) |
| KN-775 (endometrial, post-Pt) | fragile, post-platinum, mixed MMR | 30% | ~0.55 (highest) |
| LEAP-002 (HCC CP-A) | cirrhotic, BCLC B-C | 26% | ~0.10 (collapsed synergy) |
| LEAP-012 (HCC + TACE) | fit + locoregional priming | ~46% | ~0.30 (TACE boosts immune context) |
3.8 Scorecard
| Dimension | Predicted | Observed | Verdict |
|---|---|---|---|
| ORR | 65–75% | 71.0% | ✓ excellent |
| CR rate | 10–18% | 16.1% | ✓ excellent |
| mPFS | 20–25 mo | 23.9 mo | ✓ excellent |
| mOS (with crossover) | ~50 mo | 53.7 mo | ✓ excellent |
| OS HR | 0.85–1.00 (NS) | 0.79 (NS) | ✓ excellent |
| Median DoR | 20–28 mo | 26.7 mo | ✓ excellent |
| Depth of response (≥50% shrinkage) | 50–60% | ~58% | ✓ excellent |
| Time to response | 1.5–2.5 mo | 1.9 mo | ✓ excellent |
| Subsequent-therapy modeling | not in v1.0 model | 83% sunitinib pts received subsequent IO | ⚠ added in v2.1 |
| Synergy coef portability across indications | assumed transferable | varies 5× (RCC vs HCC) | ✗ population-dependent |
Cross-Trial Synthesis
Prediction methodology validation matrix
| Drug / Trial | Class accuracy (ORR) | PFS accuracy | OS accuracy |
|---|---|---|---|
| Pembro KN-006 (melanoma) | 95% | match | match |
| Pembro KN-024 (NSCLC PD-L1≥50) | 90% | match | slight ↑ |
| Pembro KN-240 (HCC 2L) | 95% | match | match (signif. missed) |
| Pembro KN-204 (Hodgkin) | 95% | match | match |
| Lenva SELECT (DTC) | 95% | match | — (NR) |
| Lenva REFLECT (HCC) | 90% | match | match (NI) |
| Lenva Study 111 (endometrial) | 95% | match | match |
| Combo CLEAR (RCC) | 95% | match | match (with crossover model) |
| Combo KN-775 (endometrial) | 90% | match | match |
| Combo LEAP-002 (HCC) — negative | over-predicted by 10–15% | slight ↑ | ✗ predicted positive, observed negative |
Three insights
Universal modifiers
Synergy enhancers
• Inflamed tumor microenvironment (high CD8, high IFN-γ signature)
• Biomarker enrichment (PD-L1≥10, MSI-H, TMB-high) — adds 10–15% ORR
• Sarcomatoid features (RCC) — adds 5–10% ORR for IO+TKI
• Locoregional priming (TACE, SBRT) — relieves tumor antigen sequestration
• IO-naïve population (no prior PD-1/PD-L1 exposure)
Synergy killers
• Cirrhotic immune dysfunction (HCC) — collapses IO contribution
• Heavy prior cytotoxic exposure — depletes immune repertoire
• ECOG ≥2 → reduced ORR ×0.75, PFS ×0.80
• Active autoimmune disease on immunosuppression
• Liver-only metastatic burden — lower IO response in some series
External Validation — Efficacy Predictions vs 3 Independent Trials
The framework — calibrated against CLEAR (1L RCC, s = 0.45) — is now validated against three independent combination trials representing different population fitness, dose levels, and disease contexts: KEYNOTE-775 (endometrial, post-Pt — synergy preserved), LEAP-002 (HCC, cirrhotic — negative trial, synergy collapsed), LEAP-012 (HCC + TACE — locoregional rescue restored partial synergy). Each is treated as a held-out cohort. Predicted values use the indication-specific synergy coefficient and population fitness modifiers — no per-trial efficacy tuning.
CLEAR (calibration cohort) — sanity check
Input profile
Lenva 20 mg + pembro 200 mg Q3W · age <65 (median 62) · ECOG 0-1 · IO-naïve · ccRCC · Synergy 0.45 · Fitness mod 1.00×
| Metric | Predicted | Observed (Motzer 2021/2024) | Δ | Note |
|---|---|---|---|---|
| ORR | 69% | 71.0% | +2% match | core synergy model anchor |
| CR rate | 14% | 16.1% | +2% match | — |
| mPFS | 22.5 mo | 23.9 mo | +1.4 mo match | — |
| mOS | 50 mo | 53.7 mo | +3.7 mo match | crossover-adjusted prediction |
| OS HR | 0.85 | 0.79 (NS) | match | — |
KEYNOTE-775 / Study 309 — Endometrial 2L post-platinum
Input profile (Makker 2022, N=411 combo arm)
Lenva 20 mg + pembro 200 mg Q3W · age median 65 (use 65–74 bracket: ×0.93 fitness) · ECOG 0-1 · prior cytotoxic 2+ lines (×0.85 ORR mod) · pelvic RT in ~30% · Synergy 0.55 (highest — both arms weak alone) · Fitness mod 0.79×
| Metric | Predicted | Observed (Makker 2022) | Δ | Note |
|---|---|---|---|---|
| ORR (all-comer) | 32% | 30.3% | −1.7% match | fitness mod calibration validates on the down-side |
| ORR (pMMR subset) | 30% | 30.0% | match | — |
| CR rate | 5–8% | 5.4% | match | — |
| mPFS (all-comer) | 6.5 mo | 6.6 mo | +0.1 mo match | vs chemo 3.8 mo (HR 0.56) |
| mOS | 17 mo | 18.3 mo | +1.3 mo match | vs chemo 11.4 (HR 0.62) — positive primary endpoint |
| Median DoR | 10–14 mo | 14.4 mo | +0–4 mo match | — |
LEAP-002 — HCC 1L, cirrhotic Child-Pugh A — NEGATIVE TRIAL
Input profile (Llovet 2023, N=395 combo arm)
Lenva 12 mg (≥60 kg) / 8 mg (<60 kg) + pembro 200 mg Q3W · age 65 · ECOG 0-1 · cirrhosis CP-A · Synergy v1.0 = 0.30 (over-predicted) · Synergy v2.0 = 0.10 (learned from negative result) · Fitness mod 0.92×
| Metric | Predicted v1.0 | Predicted v2.0 | Observed (Llovet 2023) | Δ (v2.0) |
|---|---|---|---|---|
| ORR | 32% | 27% | 26.1% | −0.9% match |
| mPFS | 10.5 mo | 8.4 mo | 8.2 mo | −0.2 mo match |
| mOS | 23.5 mo | 21.0 mo | 21.2 mo | +0.2 mo match |
| OS HR (vs lenva mono control) | 0.74 (positive) | 0.85 (NS) | 0.84 (NS) | match — predicted negative |
| Trial conclusion | predicted positive ✗ | predicted negative ✓ | negative | only v2.0 calls it correctly |
LEAP-012 — HCC + TACE (locoregional + systemic)
Input profile (Llovet 2024, combo arm N≈241)
Lenva 12 mg + pembro 400 mg Q6W + TACE (~3 sessions over 6 mo) · age 65 · ECOG 0 · cirrhosis CP-A · TACE input: synergy boost 0.10 → 0.30 (locoregional priming) · Fitness mod 0.92×
| Metric | Predicted v1.0 (no TACE input) | Predicted v2.0 (TACE input) | Observed (Llovet 2024) | Δ (v2.0) |
|---|---|---|---|---|
| ORR | 27% | 43% | ~46% | +3% match |
| mPFS | 8.4 mo | 13.5 mo | 14.6 mo | +1.1 mo match |
| PFS HR vs TACE+placebo | 0.85 | 0.65 | 0.66 | match |
| mOS | 21 mo | ~30 mo (preliminary) | NR (interim) | awaiting final |
| Trial conclusion | predicted negative | predicted positive (PFS) | positive PFS | v2.0 correct |
Cross-trial validation summary
| Trial | ORR Δ (v2.0) | mPFS Δ (v2.0) | mOS Δ (v2.0) | Synergy coef predicted correctly? |
|---|---|---|---|---|
| CLEAR (RCC, fit) | +2% | +1.4 mo | +3.7 mo (with crossover) | ✓ calibration |
| KN-775 (endometrial, fragile) | −1.7% | +0.1 mo | +1.3 mo | ✓ s = 0.55 validates highest synergy at lowest baselines |
| LEAP-002 (HCC, CP-A) — negative | −0.9% (v2.0) | −0.2 mo (v2.0) | +0.2 mo (v2.0) | ✗ v1.0 over-predicted; ✓ v2.0 calibrated to s = 0.10 |
| LEAP-012 (HCC+TACE) | +3% (v2.0 with TACE) | +1.1 mo | awaiting OS readout | ✓ TACE input restores synergy from 0.10 → 0.30 |
What works (across all 4 trials)
✓ ORR predicted within 5% in all 4 trials post-v2.0
✓ mPFS predicted within 2 mo in all 4 trials post-v2.0
✓ Max-of-arms × synergy formula validates across positive and negative outcomes
✓ Cirrhotic immune dampening identified post-LEAP-002, calibrated to s = 0.10
✓ Locoregional priming (TACE) input rescues synergy from 0.10 → 0.30 (LEAP-012)
✓ Cross-confirmation with safety analysis: KN-775's high efficacy benefit + high fatal risk travel together
What breaks (model gaps identified)
✗ v1.0 synergy coefficients were not transferable across indications (RCC → HCC failure)
✗ OS prediction without crossover/subsequent-therapy modeling is unreliable
✗ Median PFS is a poor IO endpoint — durability tail is what matters; consider 24-mo PFS rate instead
✗ Biomarker enrichment (PD-L1 CPS, MSI-H, TMB) was not in v1.0 — added in v2.1
✗ Subsequent immunotherapy in control arm collapses OS HR — population-specific subsequent-therapy modeling needed
Proposed model refinements (v2.0 + v2.1)
| Refinement | Trigger / Source | Mechanism | Expected impact |
|---|---|---|---|
| Indication-specific synergy coefficients | LEAP-002 negative result | RCC s = 0.45 · Endometrial s = 0.55 · HCC s = 0.10 · HCC+locoregional s = 0.30 | Reproduces all 4 trials within 5% ORR |
| Cirrhotic immune-dampening flag | LEAP-002 vs CLEAR contrast | CP-A: synergy ×0.30 (relative to indication baseline); CP-B/C: synergy ×0.10 | HCC ORR: 32% → 27% |
| Locoregional priming input | LEAP-012 out-of-scope in v1.0 | TACE/SBRT/RFA → adds 0.20 to synergy coef in HCC; 0.10 in non-HCC | HCC+TACE PFS: 8.4 → 13.5 mo |
| Biomarker enrichment | KN-024 vs KN-042 contrast | PD-L1 high / MSI-H / TMB-high → ORR ×1.4, PFS ×1.5 | Aligns predictions with biomarker-stratified subgroups |
| Subsequent-therapy OS dilution | CLEAR mOS observation | If control arm allows IO crossover → predicted OS HR ×1.3 (toward null) | CLEAR mOS HR: 0.55 (raw) → 0.85 (post-crossover) |
| 24-mo PFS rate alongside mPFS | IO efficacy patterns (KN-045 etc.) | Median PFS hides durability tail; report 24-mo PFS rate as second metric | Reframes IO efficacy in terms of curable subset |
Real-World Calibration Layer v2.3
Clinical trials systematically over-represent efficacy. Selection criteria (ECOG 0–1, intact organ function, age ceilings, comorbidity exclusions) enrich for fitness; protocolized imaging cadence detects responses faster and more completely; structured follow-up captures durability that real-world practice misses (lost to follow-up, transitions to hospice). The result: published trial efficacy is systematically optimistic when applied to community-oncology patients.
The calculator includes a Setting toggle (Clinical Trial / Real-World Practice) that applies literature-derived attenuation factors when switched to real-world mode. The factors are summarized below.
Adjustment coefficients (Trial → Real-World)
| Dimension | Multiplier | Direction | Mechanism |
|---|---|---|---|
| ORR | ×0.85 | ↓ ~15% | Non-protocolized imaging cadence, mixed RECIST adherence, broader baseline disease burden, less aggressive premedication |
| Median PFS | ×0.85 | ↓ ~15% | Compounding effect of lower response rates + earlier progression detection threshold variability |
| Median OS | ×0.85 | ↓ ~15% | Patient-mix related (more advanced/refractory disease, more comorbidity-related deaths, less access to subsequent therapy) |
| CR rate | ×0.70 | ↓ ~30% | CR is the metric most sensitive to imaging cadence and confirmation requirements |
| Median DoR | ×0.80 | ↓ ~20% | Less intensive surveillance → later detection of progression in responders → DoR appears truncated |
| 2-yr PFS rate | ×0.75 | ↓ ~25% | Tail of the curve is most sensitive to dropout and crossover |
Supporting literature
| Study type / source | Indication | Key finding vs trial |
|---|---|---|
| Flatiron Health real-world RCC (Bilen 2023, ASCO GU) | RCC, lenva+pembro | Median age 67 vs 62 in CLEAR; ECOG 2+ in 18% vs <5%; mPFS 14.2 mo vs CLEAR 23.9 mo (~40% attenuation in unselected pop) — exceeds the −15% calculator estimate |
| Hatakeyama 2024 (Japanese RW) | RCC, lenva+pembro | ORR 53% real-world vs 71% CLEAR; mPFS 16 mo vs 23.9; consistent with multiplicative attenuation |
| Adra 2023 (real-world endometrial) | Endometrial, lenva+pembro | ORR 23% real-world vs 30% KN-775 (~25% relative attenuation) — confirms ×0.85 efficacy mult is a floor, not a ceiling |
| Pinato 2022 (IMbrave150 RWE) | HCC, atezo+bev | Trial-to-RWE ORR ratio ~0.80; mOS ratio ~0.80 — supports cross-combo applicability of multipliers |
| Khaki 2021 IO meta-RWE | Pembrolizumab mono, multiple | RW ORR ~0.85 trial; mOS ratio 0.80–0.90 across indications |
| Bossi 2023 lenvatinib RW | DTC, HCC | RW ORR 56% vs SELECT 65% (−14%); dose interruptions occur faster (median 4 wk vs 6 wk) |
| NCDB / Medicare claims | Various | OS attenuation in unselected populations vs trial typically 10–20% for IO regimens; greater for elderly subsets and ECOG 2+ |
Implications
For clinical practice
When counseling patients in community oncology, quote the real-world numbers, not the trial numbers. A patient told "71% response rate" based on CLEAR who instead has a 53% chance has been mis-counseled.
Conversely, in tertiary academic centers with intensive imaging and supportive infrastructure, trial numbers may be appropriate.
Pair efficacy and safety counseling using both calculators in real-world mode — the resulting risk/benefit ratio differs meaningfully from the trial-mode pairing.
For drug development & regulatory
Trial enrollment criteria significantly compress both safety and efficacy envelopes — but in opposite directions. Safety looks better than reality; efficacy also looks better than reality. The risk/benefit ratio observed in trials is therefore neither over- nor under-stated; it shifts but the relative shape may be preserved.
Pragmatic Phase IV trials and mandatory post-approval RWE studies are the only way to close both gaps simultaneously. The framework's RWE adjustment is a stop-gap until trial-vs-RWE deltas are reported routinely for both efficacy and safety endpoints.
Response Onset Timeline — Combination Regimen
Typical onset and peak windows for response, depth amplification, and durability decisions on lenva 20 mg + pembro 200 mg Q3W. Light bar = possible window; solid bar = peak / typical window; black tick = median time. X-axis: 0–4 weeks (high resolution), 4–12, 12–26, 26–52+ weeks.
Response Pattern Decision Trees — Atypical IO+TKI Patterns
For ambiguous response patterns in the combination regimen. Each card walks through the clinical algorithm for distinguishing immune-mediated patterns (pseudoprogression, mixed response, late response, hyperprogression, durable response after discontinuation) from conventional response/progression. Misclassification of pseudoprogression alone causes ~2% premature treatment discontinuation in IO+TKI cohorts — getting the pattern right preserves ongoing benefit.
Pseudoprogression 3–7% of patients
Apparent radiographic progression (≥20% growth, new lesion) followed by subsequent regression on continued therapy. Driven by inflammatory infiltrate within tumor, transient lesion enlargement before immune-mediated kill. Reported in 3–7% of IO+TKI patients (lower than IO mono ~5–10% because TKI provides direct cytostatic counterbalance).
Step 1 — Clinical status at apparent progression
Step 2 — Imaging pattern
Step 3 — Confirmatory rescan (4–8 weeks later)
Mixed response (heterogeneous lesion behavior) ~10–15%
Some lesions shrink while others grow on the same scan. Reflects intra-patient tumor heterogeneity (clonal divergence, microenvironment differences). Common in IO+TKI because the two arms reach different metastatic sites with different efficacy.
Step 1 — Quantify the mix
Step 2 — Site-specific patterns
Late / delayed response (response after first apparent SD/PD) ~3% of long responders
Tumor shrinkage that begins after week 12–16 — outside the typical TKI-driven early response window. Driven by delayed IO immune response (T-cell repertoire expansion takes months in some patients).
Step 1 — Pattern recognition
Hyperprogression (paradoxical acceleration on IO) ~1–4% rare
Tumor growth rate (TGR) at least doubles compared to pre-treatment baseline within the first 8 weeks of IO. Mechanism debated — possibly Treg-driven enhancement, MDM2 amplification, EGFR mutation. Less common in IO+TKI than IO mono because TKI provides countervailing cytostatic signal.
Step 1 — Quantify TGR change
Step 2 — Risk factors (helpful but not deterministic)
Durable response after discontinuation ~30–40% of CRs
Continued response (CR or sustained PR) after planned or unplanned discontinuation of the regimen. Particularly common with IO mono (~40% of CRs sustain off-treatment); less established for IO+TKI but emerging data suggests ~30%.
Step 1 — Reason for discontinuation
Step 2 — Surveillance schedule for off-treatment CR
Key References
Pembrolizumab monotherapy — efficacy
- Robert C et al. NEJM 2015;372:2521–32 — KEYNOTE-006 (melanoma, ORR/PFS/OS)
- Reck M et al. NEJM 2016;375:1823–33; Brahmer J JCO 2023 — KEYNOTE-024 5-yr update
- Mok TSK et al. Lancet 2019;393:1819–30 — KEYNOTE-042
- Bellmunt J et al. NEJM 2017;376:1015–26; Fradet Y Ann Oncol 2019 — KEYNOTE-045
- Kuruvilla J et al. Lancet Oncol 2021;22:512–24 — KEYNOTE-204 Hodgkin
- Marabelle A et al. J Clin Oncol 2020;38:1–10 — KEYNOTE-158 MSI-H
- Finn RS et al. J Clin Oncol 2020;38:193–202 — KEYNOTE-240 HCC 2L
- Balar AV et al. Lancet Oncol 2017;18:1483–92 — KEYNOTE-052 urothelial 1L cis-ineligible
- McDermott DF et al. Lancet Oncol 2021 — KEYNOTE-427 (pembro mono RCC, used as combo input)
Lenvatinib monotherapy — efficacy
- Schlumberger M et al. NEJM 2015;372:621–30 — SELECT (DTC ORR 64.8%, mPFS 18.3 mo)
- Kudo M et al. Lancet 2018;391:1163–73 — REFLECT (HCC, vs sorafenib NI)
- Vergote I et al. J Clin Oncol 2020;38:2961–8 — Study 111 endometrial mono
- Brose MS et al. Lancet Oncol 2022;23:1153–64 — Study 211 (DTC dose 18 vs 24)
- Sato J et al. Lancet Oncol 2020;21:843–50 — REMORA thymic carcinoma
Combination — CLEAR (calibration cohort)
- Motzer R et al. NEJM 2021;384:1289–300 — CLEAR primary (NCT02811861); ORR 71.0%, mPFS 23.9 mo, mOS 53.7 mo
- Motzer R et al. J Clin Oncol 2024;42:1222–8 — CLEAR final OS analysis (HR 0.79 NS)
- Grünwald V et al. Lancet Oncol 2022;23:768–80 — CLEAR HRQoL
Comparator combinations (validation cohorts)
- Makker V et al. NEJM 2022;386:437–48 — KEYNOTE-775 (endometrial 2L; ORR 30%, mPFS 6.6 mo, mOS 18.3 mo positive)
- Llovet JM et al. J Clin Oncol 2023;41:1796–807 — LEAP-002 HCC (negative; ORR 26.1%, OS HR 0.84 NS)
- Llovet JM et al. Lancet 2024 — LEAP-012 (HCC + TACE; PFS HR 0.66 positive interim)
- Choueiri TK et al. NEJM 2021;384:829–41 — CheckMate-9ER (nivo+cabo RCC)
- Powles T et al. NEJM 2019;381:830–9; Plimack E Ann Oncol 2024 — KEYNOTE-426 (pembro+axi RCC)
- Finn RS et al. NEJM 2020;382:1894–905 — IMbrave150 (atezo+bev HCC; only positive HCC IO+anti-VEGF)
- Eskander RN et al. NEJM 2023;388:2159–70 — NRG-GY018 (pembro+chemo endometrial 1L)
Real-world efficacy literature
- Bilen MA et al. ASCO GU 2023 — Flatiron Health real-world lenva+pembro RCC (mPFS 14.2 mo)
- Hatakeyama S et al. 2024 — Japanese real-world lenva+pembro RCC
- Adra N et al. 2023 — Real-world endometrial lenva+pembro
- Pinato DJ et al. 2022 — IMbrave150 RWE
- Khaki AR et al. 2021 — Pembrolizumab mono real-world meta-analysis
- Bossi P et al. 2023 — Lenvatinib real-world DTC and HCC
All trial efficacy values as published. Class predictions are author-derived from mechanistic baselines. Synergy coefficients calibrated against CLEAR and refined post-hoc against KN-775, LEAP-002, LEAP-012.