Methodology · M1

IBC Reasoning Audit.

19 agents review every Atlas Bio output before it reaches a clinician or reviewer. 11 core agents handle the analytical pipeline; 8 reasoning nodes audit the result through eight independent lenses. Every prediction passes through all eight.

R01 through R08.

NodeDomainCatches
R01 · CausalCausal-chain integrity"Correlation reported as causation" — model output that implies a causal claim its training data can't support
R02 · ContradictionInternal consistencyOutput that contradicts an earlier claim in the same report, or contradicts a published reference
R03 · ConfidenceCalibration of stated uncertaintyOver-confidence on thin evidence; under-confidence when evidence is strong
R04 · MechanisticBiology / chemistry plausibilityOutputs that ignore well-established mechanism (e.g., predicting CNS penetration with no transcytosis story)
R05 · AnalogicalCross-case comparisonOutputs that ignore an obvious analog (e.g., predicting IO+TKI synergy in HCC without referencing LEAP-002)
R06 · AbductiveInference to the best explanationOutput that picks one hypothesis without considering alternatives
R07 · TemporalTime-ordering integrityHindsight contamination — using future data to "predict" past events
R08 · MetaAudit of the auditWhen the other 7 nodes disagree, R08 arbitrates and surfaces the disagreement to the human reviewer

Six-step cycle.

  1. Ingest. Analytical pipeline produces a draft prediction (with calibration, references, mechanism notes).
  2. R01–R07 run in parallel. Each node returns a verdict (pass / soft flag / hard flag) and an explanation.
  3. R08 aggregates. If all 7 pass → ship. If hard flag → block and surface the failure. If soft flags → annotate the output with caveats.
  4. Human review. Reviewer can override an R08 verdict with a comment that becomes part of the audit record.
  5. Re-audit. If the prediction is revised, the full 19-agent pipeline re-runs on the revised version.
  6. Lock. Final version SHA-256 hashed and added to the pre-registration ledger (see M2).

Real failures the audit found.

Alpha Predictor hindsight contamination (R02 + R07)

An early version reported 96.6% blind-test accuracy on Q4 2024 earnings catalysts. R02 (Contradiction) and R07 (Temporal) flagged the result: 5 of 14 trades had been added to the test set after their outcomes were known. Honest re-test after stripping: 91%. We published the failure and corrected the model. Read the post-mortem →

LEAP-002 synergy over-prediction (R04 + R05)

Pre-trial predictions used the RCC-calibrated synergy coefficient (s = 0.30) transferred to HCC. R04 (Mechanistic) flagged that cirrhotic immune dysfunction would likely collapse IO contribution. R05 (Analogical) flagged that KEYNOTE-240 (pembro mono HCC) barely cleared significance — adding a TKI wouldn't reliably push past. Both flags were soft. The trial was negative. Post-hoc, we recalibrated to s = 0.10 and the v2.0 calculator now reproduces the negative outcome.

KN-775 Gr3 under-prediction (R04)

v1.0 predicted Gr3 TRAE 70.3% vs observed 88.9%. R04 (Mechanistic) flagged that post-platinum + pelvic-RT populations have substrate vulnerability that shifts Gr1-2 events into Gr3 territory. v2.0 added the substrate Gr3-shift coefficient (×1.25) and the prediction now matches.

Source files.

Each reasoning node has its own prime directive as a styled HTML file in agents/IBC_Reasoning/ — 8 .md + 8 .html files. The prime directives are reviewable by clinicians and regulators; they describe what each node does, its evidence base, and known limitations. Request access →

Need an independent audit of your model?

The 19-agent pipeline can be applied to your model's outputs in a single briefing. Useful for internal model review, regulatory submissions, or board-level second opinions.

Contact Us →