R01 through R08.
| Node | Domain | Catches |
|---|---|---|
| R01 · Causal | Causal-chain integrity | "Correlation reported as causation" — model output that implies a causal claim its training data can't support |
| R02 · Contradiction | Internal consistency | Output that contradicts an earlier claim in the same report, or contradicts a published reference |
| R03 · Confidence | Calibration of stated uncertainty | Over-confidence on thin evidence; under-confidence when evidence is strong |
| R04 · Mechanistic | Biology / chemistry plausibility | Outputs that ignore well-established mechanism (e.g., predicting CNS penetration with no transcytosis story) |
| R05 · Analogical | Cross-case comparison | Outputs that ignore an obvious analog (e.g., predicting IO+TKI synergy in HCC without referencing LEAP-002) |
| R06 · Abductive | Inference to the best explanation | Output that picks one hypothesis without considering alternatives |
| R07 · Temporal | Time-ordering integrity | Hindsight contamination — using future data to "predict" past events |
| R08 · Meta | Audit of the audit | When the other 7 nodes disagree, R08 arbitrates and surfaces the disagreement to the human reviewer |
Six-step cycle.
- Ingest. Analytical pipeline produces a draft prediction (with calibration, references, mechanism notes).
- R01–R07 run in parallel. Each node returns a verdict (pass / soft flag / hard flag) and an explanation.
- R08 aggregates. If all 7 pass → ship. If hard flag → block and surface the failure. If soft flags → annotate the output with caveats.
- Human review. Reviewer can override an R08 verdict with a comment that becomes part of the audit record.
- Re-audit. If the prediction is revised, the full 19-agent pipeline re-runs on the revised version.
- Lock. Final version SHA-256 hashed and added to the pre-registration ledger (see M2).
Real failures the audit found.
Alpha Predictor hindsight contamination (R02 + R07)
An early version reported 96.6% blind-test accuracy on Q4 2024 earnings catalysts. R02 (Contradiction) and R07 (Temporal) flagged the result: 5 of 14 trades had been added to the test set after their outcomes were known. Honest re-test after stripping: 91%. We published the failure and corrected the model. Read the post-mortem →
LEAP-002 synergy over-prediction (R04 + R05)
Pre-trial predictions used the RCC-calibrated synergy coefficient (s = 0.30) transferred to HCC. R04 (Mechanistic) flagged that cirrhotic immune dysfunction would likely collapse IO contribution. R05 (Analogical) flagged that KEYNOTE-240 (pembro mono HCC) barely cleared significance — adding a TKI wouldn't reliably push past. Both flags were soft. The trial was negative. Post-hoc, we recalibrated to s = 0.10 and the v2.0 calculator now reproduces the negative outcome.
KN-775 Gr3 under-prediction (R04)
v1.0 predicted Gr3 TRAE 70.3% vs observed 88.9%. R04 (Mechanistic) flagged that post-platinum + pelvic-RT populations have substrate vulnerability that shifts Gr1-2 events into Gr3 territory. v2.0 added the substrate Gr3-shift coefficient (×1.25) and the prediction now matches.
Source files.
Each reasoning node has its own prime directive as a styled HTML file in agents/IBC_Reasoning/ — 8 .md + 8 .html files. The prime directives are reviewable by clinicians and regulators; they describe what each node does, its evidence base, and known limitations. Request access →
Need an independent audit of your model?
The 19-agent pipeline can be applied to your model's outputs in a single briefing. Useful for internal model review, regulatory submissions, or board-level second opinions.
Contact Us →