Why eight levels.
A common failure mode in capsid prediction is mixing evidence at different abstraction layers and over-claiming. A pocket-detection algorithm (physics) and a clinical-durability extrapolation (outcome) are not the same kind of prediction and shouldn't be reported with the same confidence. The L1–L8 framework forces predictions to be tagged with their evidence layer and validated against that layer's appropriate gold standard.
From physics to outcomes.
| Level | Domain | Module status | Gold standard |
|---|---|---|---|
| L1 | Physics — pocket detection (BFS), surface curvature (cotangent Laplacian), hydrophobic regions (BFS flood-fill), mutation suggestions (Kyte-Doolittle), dihedral distances | 5 of 5 DONE | PDB structures (n=5) |
| L2 | Scaffolding — surfaceome calibration (HPA nTPM), perturbation data ingestion, binding affinities | 4 core + 3 data modules DONE | 30 perturbation studies + 7 Kd |
| L3 | Receptor — receptor engagement, CPD discovery (AAVR2, Su 2025), GPR108-independent pathway (Dudek 2020) | DONE | SwissProt + published binding |
| L4 | Variant fitness — CapsidFitnessCNN (256K params) trained on 2,014 Bryant variants + ESM-2 embeddings | DONE | Bryant 2014 DMS · contrastive AUC 0.793 |
| L5 | Tropism — tissue preference mapping, receptor expectation tables (multi-tissue) | DONE | HPA + published biodist |
| L6 | Immune context — seroprevalence (corrected per Boutin 2010), neutralizing antibody escape | PARTIAL | seroprev literature meta-analysis |
| L7 | Manufacturing — yield, packaging, full/empty ratio, post-translational | DONE | vendor-disclosed yield data |
| L8 | Clinical durability — long-term transgene expression, immune-mediated loss, integration | FUTURE | NCT trials (gathering data) |
80/80 tests passing.
Each module has its own unit + integration test suite. The combined runner exercises all 8 levels in sequence against the 15-capsid calibration set. Current status: 80/80 tests passing · 72 training pairs exported for contrastive learning · zero known failures.
Tests are gated on real data only — there is no synthetic-data fallback. If a module can't reach its data source (HPA, Bryant, SwissProt, PDB), the test fails loudly rather than silently substituting a stub.
What's next.
- L6 completion — neutralizing antibody escape prediction module. Currently partial; full release Q3 2026.
- L8 first pass — clinical durability proxy model using the existing 20-trial AAV trial database. Pre-registered against an additional 5-trial holdout.
- L4 expansion — extend training set from 2,014 Bryant variants to the 800K+ published fitness variant pool identified during Session 11.
- Cross-level confidence calibration — formal Bayesian propagation across L1 → L8 so multi-level predictions surface their weakest link.
Want the full strategy report?
10-section "Predictive Level Strategy" deck available to verified clinicians and sponsors. Covers all 8 levels with gap analysis, roadmap, investment summary, and published-data source list.
Contact Us →