CapsidFitnessCNN — 256K parameters.
Single-task supervised model + contrastive auxiliary objective. Input is the VP1 sequence embedding stack; output is a tissue-conditional fitness score in [0, 1]. Trained on CPU (~30 min for the 2,014-variant set); GPU not required for inference.
Input features
- Real ESM-2 embeddings — 640-dim from
facebook/esm2_t30_150M_UR50D. Replaces the prior mock embeddings. - 20 amino-acid fractions — composition vector at the VP3 surface region.
- Isoelectric point (pI) — calculated from sequence, not predicted.
- GRAVY hydropathy index — Kyte-Doolittle.
- VP3 surface composition — separate from full-capsid composition; what's actually exposed.
- Net charge at physiological pH — relevant for HSPG binding and serum stability.
Headline metrics.
| Metric | Before fixes (v1) | After fixes (v2) | Delta |
|---|---|---|---|
| Contrastive AUC | 0.608 | 0.793 | +0.185 ↑ |
| Classification accuracy | 0.596 | 0.761 | +0.165 ↑ |
| AAV9 CNS rank | #3 (score 0.461) | #1 (score 0.662) | multi-scale fix |
Trained on 2,014 variants; ranks 15.
Training pool is the Bryant 2014 deep-mutational-scan of AAV2 (2,014 variants with viability/fitness annotations). Output ranking applies to all 15 modeled capsids (13 natural + 2 engineered 4DMT) and any new candidate sequence submitted via the inquiry pipeline.
Note: the model is calibrated for the AAV2-like serotype family. Highly divergent capsids (e.g., AAV5 — 60% sequence identity to AAV2) get a flag in the output indicating "out-of-distribution; treat with caution."
Two-step inquiry.
- Send a candidate list — VP1 sequences or program name with capsid spec — via inquiry.
- We return a ranked report within 5 business days: contrastive score, BBB score (if relevant), seroprevalence-adjusted clinical risk, comparison to the 15-capsid baseline.
Have a capsid program to rank?
The platform handles natural, engineered, ancestral, and chimeric capsids. Submission accepts up to 50 candidates per briefing.
Contact Us →