Methodology · The operating system

How an agentic platform earns trust.

Atlas Bio's methodology is the rule-set the agentic AI lives by — how the agents reason, where humans intervene, how every claim earns its evidence level, and why a prediction can't be wrong by accident without us catching it. Three sections: how the platform is wired (operating system), how we know we're not lying to ourselves (discipline), and why the math works (scale and honest failure).

🔒
Methodology details are gated. The public page describes the shape of the discipline. The full architecture diagram, audit-node taxonomy, HITL decision-rights, validation rubric, and closed-files list are available to verified clinicians, sponsors, and reviewers under MNDA.

How the platform is wired.

Three structural commitments that make Atlas Bio an agentic platform rather than a chatbot or a one-off model: a multi-tier role structure with decisional authority, a Human-in-the-Loop protocol with explicit override paths, and an evidence hierarchy that tags every claim by strength.

A1 · Architecture

Agentic role structure with named authorities

Not one model, not a swarm. A multi-tier agentic stack: scientific-direction agents, an orchestrator, specialist agents, an audit layer, and a Human-in-the-Loop board. Roles, handoffs, and authority are explicit — not implicit.

Multi-tier agentic stack · full architecture under MNDA 🔒 NDA

A2 · Human-in-the-Loop

Agents execute. Humans decide.

The HITL protocol is a formal contract: a defined set of decisions is non-delegable to agents (scope, kill-criteria enforcement, translational gates, regulatory commit). Other decisions are agent-owned with audit traces. Every advance between evidence tiers requires a signed human review.

Named human-only decisions · full decision-rights protocol under MNDA 🔒 NDA

A3 · Evidence hierarchy

Every claim is tagged L1, L2, or L3.

L1 (observational / correlational) is surveyed but not headlined. L2 (causal — mechanism in model) is the working evidence layer. L3 (thermodynamic / first-principles) earns the program's strongest publishable assertions. The level is on the claim, every time.

L1 / L2 / L3 tagging on every claim

Architecture diagram — named role pills, decision-rights flow, and the full reasoning-audit taxonomy are available under MNDA. Unlock above ↑
Decides Human-in-the-Loop Board ←→ Lead Scientist (CSO-grade)
Coordinates Program Orchestrator
Executes CP-01..CP-15 Clinical Pharm MI-01 Imaging Program-specific specialist agents
Audits 8 IBC Reasoning Nodes 11 Core Audit Agents

Active platform: 249 agents · 15 clinical-pharmacology specialists · 8 IBC reasoning nodes · 11 core audit agents · 1 Lead Scientist per program.

Lead Scientist — 5 non-delegable decisions

  1. Hypothesis prioritization. Which mechanisms get platform resources this quarter; which are deferred.
  2. Compound selection & ejection. Which compounds enter Stage 6 in vivo; which are removed from the program.
  3. Blend composition. The rationale for every multi-compound stack — signed before Pfizer-grade validation begins.
  4. Kill-criteria enforcement. When a candidate fails a pre-registered kill criterion, the Lead Scientist enforces removal regardless of sunk cost.
  5. Translational readiness gate. Advance from single_tool_only evidence to multi_tool_agreement evidence — the gate to IND assembly.
LevelClassWhat it looks like
L1Correlational / observationalEpidemiology associations, registry trends, single-cohort signals. Surveyed; never headlined.
L2Causal — mechanism in modelEffect demonstrated in an animal model, knockdown rescue, dose-response in vivo. The working evidence layer for recommendations.
L3Thermodynamic / first-principlesValidated binding ΔG, mass-balance kinetics, physically-grounded mechanism. The strongest publishable claims.

How we know we're not lying to ourselves.

Four disciplines that catch the four most common failure modes in computational therapeutics: hindsight bias, reasoning failure, over-fitting, and trial-vs-real-world divergence.

A4 · Pre-registration

Predictions locked before evaluation

Hypotheses, primary endpoints, statistical analysis plan, and kill criteria are SHA-256 hashed and time-stamped before any data collection. HARKing — hypothesizing after results are known — is treated as scientific misconduct, not a stylistic preference.

SHA-256 hashes on every Phase IIa · pre-data SAP sign-off required

A5 · Reasoning audit

Multi-node reasoning audit on every output

Every output is reviewed by a multi-node reasoning audit covering causal, contradiction, and temporal axes, among others. Each node has a published prime directive; audit traces are stored against the original prediction.

Multi-node audit · full taxonomy & node behaviors under MNDA 🔒 NDA

A6 · Cross-domain validation

One parameter set, multiple independent cases

The mark of a real model is that it works across cases without per-case tuning. A single parameter set is tested against multiple independent trials per program. Negative-trial reproduction (the model has to predict failures too) is the hardest test we routinely pass.

Cross-domain validation · negative-test reproduction · validation rubric under MNDA 🔒 NDA

A7 · RWE calibration

Live biomedical data keeps the model honest

Trial populations are systematically fitter than community populations. Continuous calibration against ClinicalTrials.gov, FDA AERS/FAERS, PubMed, EMR cohorts, patient registries, and disease-specific consortia is what makes a prediction translate from trial to clinic.

Public biomedical pipelines · EMR-cohort calibration partnerships under MNDA 🔒 NDA

IBC reasoning audit — 8 nodes

NodeAudits for
R01 · CausalConfound, reverse causation, mediator confusion
R02 · ContradictionInternal inconsistency between claims; data contradicting model
R03 · ConfidenceOver-confidence vs evidence base; mis-calibrated certainty
R04 · MechanisticPlausibility of stated mechanism vs known biology / physics
R05 · AnalogicalMis-applied analogy across modalities, species, or indications
R06 · AbductiveBest-explanation reasoning; competing hypothesis enumeration
R07 · TemporalTime-order violations, future-data leakage, hindsight contamination
R08 · Meta-reasoningReasoning about the audit itself; cross-node consistency
Failure modeHow it shows upAtlas Bio defense
Hindsight contamination"Predicting" outcomes that were already known when the model was tuned.A4 · SHA-256 pre-registration before the validation event.
Reasoning failuresPlausible-looking output that fails a basic causal, mechanistic, or contradictory check.A5 · multi-node reasoning audit on every output.
Over-fittingModel excels on training case but breaks on the next one.A6 · cross-domain validation with a single parameter set; negative-test reproduction.
Trial vs RWE divergenceTrial population is fitter than community; predictions are systematically optimistic.A7 · RWE calibration multipliers from live biomedical data sources.
Black-box opacity"Trust the model" without explanation; clinicians and regulators reject.Calculator + comparator + decision trees expose every input → output. Plus L1/L2/L3 evidence tagging on every claim.

Scale where it matters, and the discipline to fail in public.

The agentic methodology isn't only about how predictions are checked. It's about which checks become possible at agent speed — and what happens when the audit catches the firm itself.

A8 · Calibrated predictions at agent speed.

Modern pre-clinical screening is no longer bottlenecked on compound count — it's bottlenecked on whether the predictions translate to humans. The agentic platform attacks that second bottleneck: agents run the breadth (literature, RWE, public catalogs, in-silico scoring) continuously, humans focus the irreplaceable judgment work on a calibrated shortlist. The result is fewer candidates entering expensive wet-lab — but each with a measured prior probability of working.

104–106Candidate space the in-silico stack handles per program
< 50Compounds typically advanced per round to wet-lab
106+Agents in the active platform
A8b · Honest failure

The audit caught the firm itself.

An internal Atlas Bio prediction engine once reported 96.6% blind-test accuracy on a Q4 2024 catalyst set. The audit's reasoning nodes flagged the result — finding that several entries had been added after their outcomes were known. The "blind" test wasn't blind.

Honest re-test after stripping the contaminated entries: 91%. Both numbers are preserved on the pre-registration ledger. The contamination is documented as the anchor failure case for a forthcoming manuscript on the reasoning-audit pipeline.

96.6% → 91% · contamination caught internally · published, not buried

Closed files

Theses formally closed.

Five theses have been formally closed based on 2024–2026 evidence. Closing a thesis isn't an embargo — it's a published judgment that reopening requires L2-or-better data, not opinion. Closed files prevent the platform from rediscovering its own dead ends and prevent partners from being sold ideas the field has already ruled out.

5 closed files in the Longevity program · full list & rationale on request

Want the methodology pack?

Briefings can be focused on methodology alone — the agentic architecture diagram, the HITL protocol, the audit-node taxonomy, the validation rubric, and the RWE infrastructure. Useful for regulators, partner reviewers, and skeptical scientific advisors.

Contact Us →

🔒 Unlock full methodology — NDA required

The full methodology pack (architecture diagram, reasoning-node taxonomy, HITL decision-rights, closed files) is shared under a mutual non-disclosure agreement. Acknowledge the terms below and unlock the gated sections.

Mutual NDA — methodology disclosure

By unlocking the gated methodology content on this page, the recipient agrees that the disclosed materials constitute the confidential proprietary methodology of Atlas Bio. The recipient shall:

1. Hold the disclosed materials in strict confidence and not reproduce, distribute, or disclose them to any third party.
2. Use the materials solely for the purpose of evaluating Atlas Bio's platform for a potential clinical, research, or sponsor engagement.
3. Not reverse-engineer, replicate, or use the materials to compete with Atlas Bio.
4. Return or destroy all materials upon request.

This site-level NDA is preliminary and operates as an interim acknowledgment. A formal MNDA may be required for ongoing engagement. Atlas Bio reserves the right to log the recipient's name, organization, and access timestamp. Site marked no-index · shared by invitation only.