What the models actually learn from
Capsid machine learning rests on measured libraries. In the most widely cited example, researchers applied deep learning to a 28-amino-acid segment of the AAV2 capsid protein and generated 201,426 variants, yielding 110,689 viable engineered capsids, 57,348 of which surpass the average diversity of natural AAV serotype sequences (Bryant et al., Nature Biotechnology 2021).
That is the shape of the data: large assayed variant sets over a narrow region of one parent capsid, plus structural files, published binding measurements and biodistribution reports. A model is only as transferable as that evidence base, which is why a capsid prediction should always arrive with the data it was trained on named.
- Deep mutational scan libraries with viability or fitness labels
- Protein language model embeddings of the capsid sequence
- Physicochemical features: amino-acid composition, isoelectric point, hydropathy, net charge at physiological pH
- Receptor sequences and published binding affinity measurements
- Reported in vivo biodistribution from peer-reviewed rodent and primate studies
Two different jobs: generation and ranking
Generative capsid design proposes new sequences inside a region of sequence space the model believes is viable. Discriminative ranking takes a candidate list you already have — natural serotypes, engineered variants, ancestral reconstructions — and orders it for a specific tissue question.
For most sponsors the second job is the one with budget attached. The question is not 'can we invent a capsid' but 'which of these twenty do we put into animals first'. Atlas Bio's published positioning is explicit about this: its AAV platform is built for pre-NHP ranking, translation-trap detection and audit-traceable evidence aggregation, and it states plainly that it does not predict non-human-primate outcomes and does not substitute for in vivo validation.
The failure modes that matter
Atlas Bio's public AAV page names eight open problems in the field, and most of them are generalization problems rather than modelling problems. Models trained mostly on natural serotypes mispredict engineered variants carrying insertions the training set never saw. Public binding datasets are dominated by a few receptor classes, so minority classes get silently mispredicted. Receptor identity itself is contested for several clinically important capsids, and a single hard label encodes that ambiguity as false confidence.
The most expensive failure is species specificity. A headline accuracy figure measured on a held-out split of the same library says very little about an engineered candidate from a different clade. Honest reporting keeps the in-sample number and the out-of-distribution number side by side, and attaches a species flag to any central-nervous-system claim.
What a usable output looks like
A capsid prediction that a scientific reviewer can act on carries four things: a ranked list, the named factors that drove each score, an uncertainty bound, and the limitations that apply to that specific candidate — including an out-of-distribution flag when the sequence is far from the training family.
Atlas Bio publishes performance for its fitness module in exactly that paired form, reporting 'Contrastive AUC 0.793 · accuracy 0.761' for the current version against 0.608 and 0.596 for the prior one, with the changes that produced the difference described. Numbers reported this way can be argued with. Numbers reported alone cannot.
Atlas Bio's Capsid Intelligence platform ranks natural, engineered, ancestral and chimeric capsids by tissue-specific fitness with per-candidate uncertainty, named contributing factors and out-of-distribution flags, under a pending United States patent (63/986,270).
Can AI replace non-human primate studies for capsids?
No. Published capsid models rank candidates and flag risks before animal work; they do not predict primate outcomes. Atlas Bio states this directly on its AAV page. The value is deciding which candidates deserve a cohort, not removing the cohort.
Why do capsid models fail on engineered variants?
Training sets are dominated by natural serotypes and deep mutational scans of one or two parents. Engineered capsids carry insertions and mutations outside that distribution, so the model projects them onto the wrong neighbourhood and reports confident, wrong scores.
What accuracy number should I ask a vendor for?
Ask for two: performance on a held-out split of the training family, and performance when an entire serotype or clade is left out. The gap between them is the number that predicts how the model behaves on your candidate.
- Bryant DH et al. Deep diversification of an AAV capsid protein by machine learning. Nature Biotechnology 2021 (PMID 33574611)
- Hordeaux J et al. The Neurotropic Properties of AAV-PHP.B Are Limited to C57BL/6J Mice. Molecular Therapy 2018 (PMC5911151)
- Hordeaux J et al. The GPI-Linked Protein LY6A Drives AAV-PHP.B Transport across the Blood-Brain Barrier. Molecular Therapy 2019 (PMC6520463)
Talk to Atlas Bio.
Sponsors, researchers and clinicians: tell us the decision you are trying to make and we will show what the platform can and cannot answer.
Contact Us →