How the platform is wired.
Three structural commitments that make Atlas Bio an agentic platform rather than a chatbot or a one-off model: a multi-tier role structure with decisional authority, a Human-in-the-Loop protocol with explicit override paths, and an evidence hierarchy that tags every claim by strength.
Agentic role structure with named authorities
Not one model, not a swarm. A multi-tier agentic stack: scientific-direction agents, an orchestrator, specialist agents, an audit layer, and a Human-in-the-Loop board. Roles, handoffs, and authority are explicit — not implicit.
Multi-tier agentic stack · full architecture under MNDA 🔒 NDA
Agents execute. Humans decide.
The HITL protocol is a formal contract: a defined set of decisions is non-delegable to agents (scope, kill-criteria enforcement, translational gates, regulatory commit). Other decisions are agent-owned with audit traces. Every advance between evidence tiers requires a signed human review.
Named human-only decisions · full decision-rights protocol under MNDA 🔒 NDA
Every claim is tagged L1, L2, or L3.
L1 (observational / correlational) is surveyed but not headlined. L2 (causal — mechanism in model) is the working evidence layer. L3 (thermodynamic / first-principles) earns the program's strongest publishable assertions. The level is on the claim, every time.
L1 / L2 / L3 tagging on every claim
Active platform: 249 agents · 15 clinical-pharmacology specialists · 8 IBC reasoning nodes · 11 core audit agents · 1 Lead Scientist per program.
Lead Scientist — 5 non-delegable decisions
- Hypothesis prioritization. Which mechanisms get platform resources this quarter; which are deferred.
- Compound selection & ejection. Which compounds enter Stage 6 in vivo; which are removed from the program.
- Blend composition. The rationale for every multi-compound stack — signed before Pfizer-grade validation begins.
- Kill-criteria enforcement. When a candidate fails a pre-registered kill criterion, the Lead Scientist enforces removal regardless of sunk cost.
- Translational readiness gate. Advance from single_tool_only evidence to multi_tool_agreement evidence — the gate to IND assembly.
| Level | Class | What it looks like |
|---|---|---|
| L1 | Correlational / observational | Epidemiology associations, registry trends, single-cohort signals. Surveyed; never headlined. |
| L2 | Causal — mechanism in model | Effect demonstrated in an animal model, knockdown rescue, dose-response in vivo. The working evidence layer for recommendations. |
| L3 | Thermodynamic / first-principles | Validated binding ΔG, mass-balance kinetics, physically-grounded mechanism. The strongest publishable claims. |
How we know we're not lying to ourselves.
Four disciplines that catch the four most common failure modes in computational therapeutics: hindsight bias, reasoning failure, over-fitting, and trial-vs-real-world divergence.
Predictions locked before evaluation
Hypotheses, primary endpoints, statistical analysis plan, and kill criteria are SHA-256 hashed and time-stamped before any data collection. HARKing — hypothesizing after results are known — is treated as scientific misconduct, not a stylistic preference.
SHA-256 hashes on every Phase IIa · pre-data SAP sign-off required
A5 · Reasoning auditMulti-node reasoning audit on every output
Every output is reviewed by a multi-node reasoning audit covering causal, contradiction, and temporal axes, among others. Each node has a published prime directive; audit traces are stored against the original prediction.
Multi-node audit · full taxonomy & node behaviors under MNDA 🔒 NDA
A6 · Cross-domain validationOne parameter set, multiple independent cases
The mark of a real model is that it works across cases without per-case tuning. A single parameter set is tested against multiple independent trials per program. Negative-trial reproduction (the model has to predict failures too) is the hardest test we routinely pass.
Cross-domain validation · negative-test reproduction · validation rubric under MNDA 🔒 NDA
A7 · RWE calibrationLive biomedical data keeps the model honest
Trial populations are systematically fitter than community populations. Continuous calibration against ClinicalTrials.gov, FDA AERS/FAERS, PubMed, EMR cohorts, patient registries, and disease-specific consortia is what makes a prediction translate from trial to clinic.
Public biomedical pipelines · EMR-cohort calibration partnerships under MNDA 🔒 NDA
IBC reasoning audit — 8 nodes
| Node | Audits for |
|---|---|
| R01 · Causal | Confound, reverse causation, mediator confusion |
| R02 · Contradiction | Internal inconsistency between claims; data contradicting model |
| R03 · Confidence | Over-confidence vs evidence base; mis-calibrated certainty |
| R04 · Mechanistic | Plausibility of stated mechanism vs known biology / physics |
| R05 · Analogical | Mis-applied analogy across modalities, species, or indications |
| R06 · Abductive | Best-explanation reasoning; competing hypothesis enumeration |
| R07 · Temporal | Time-order violations, future-data leakage, hindsight contamination |
| R08 · Meta-reasoning | Reasoning about the audit itself; cross-node consistency |
| Failure mode | How it shows up | Atlas Bio defense |
|---|---|---|
| Hindsight contamination | "Predicting" outcomes that were already known when the model was tuned. | A4 · SHA-256 pre-registration before the validation event. |
| Reasoning failures | Plausible-looking output that fails a basic causal, mechanistic, or contradictory check. | A5 · multi-node reasoning audit on every output. |
| Over-fitting | Model excels on training case but breaks on the next one. | A6 · cross-domain validation with a single parameter set; negative-test reproduction. |
| Trial vs RWE divergence | Trial population is fitter than community; predictions are systematically optimistic. | A7 · RWE calibration multipliers from live biomedical data sources. |
| Black-box opacity | "Trust the model" without explanation; clinicians and regulators reject. | Calculator + comparator + decision trees expose every input → output. Plus L1/L2/L3 evidence tagging on every claim. |
Scale where it matters, and the discipline to fail in public.
The agentic methodology isn't only about how predictions are checked. It's about which checks become possible at agent speed — and what happens when the audit catches the firm itself.
The audit caught the firm itself.
An internal Atlas Bio prediction engine once reported 96.6% blind-test accuracy on a Q4 2024 catalyst set. The audit's reasoning nodes flagged the result — finding that several entries had been added after their outcomes were known. The "blind" test wasn't blind.
Honest re-test after stripping the contaminated entries: 91%. Both numbers are preserved on the pre-registration ledger. The contamination is documented as the anchor failure case for a forthcoming manuscript on the reasoning-audit pipeline.
96.6% → 91% · contamination caught internally · published, not buried
Theses formally closed.
Five theses have been formally closed based on 2024–2026 evidence. Closing a thesis isn't an embargo — it's a published judgment that reopening requires L2-or-better data, not opinion. Closed files prevent the platform from rediscovering its own dead ends and prevent partners from being sold ideas the field has already ruled out.
5 closed files in the Longevity program · full list & rationale on request
Want the methodology pack?
Briefings can be focused on methodology alone — the agentic architecture diagram, the HITL protocol, the audit-node taxonomy, the validation rubric, and the RWE infrastructure. Useful for regulators, partner reviewers, and skeptical scientific advisors.
Contact Us →