The definitions that anchor the work
FDA defines real-world data as 'data relating to patient health status and/or the delivery of health care routinely collected from a variety of sources', and real-world evidence as 'the clinical evidence about the usage and potential benefits or risks of a medical product derived from analysis of RWD'. Listed sources include electronic health records, medical claims, product and disease registries, and digital health technologies.
The distinction matters commercially. Data is not evidence. A feed of claims records becomes evidence only when it is tied to a specific question with a defensible analysis, which is why FDA's framework for its real-world evidence program, required by the 21st Century Cures Act, treats methodology and data standards as the substance of the program.
Why the drift is predictable
Randomised trials are run under idealised, rigorously controlled conditions that can compromise external validity. A literature review of 52 studies across cardiology, mental health and oncology found that trial samples are highly selected and have a lower risk profile than real-world populations, with elderly patients frequently excluded (Kennedy-Martin et al., Trials 2015).
Because the selection acts in one direction, its consequences are directional too: real-world populations tend to show more high-grade toxicity and weaker effectiveness than the registration trial reported. Atlas Bio applies that as explicit calibration, publishing that trial populations are systematically fitter than community populations and that its pipelines quantify the divergence and feed multipliers into its calculators — with safety inflation and efficacy attenuation described as roughly symmetric, so the risk-benefit shape may be better preserved than either number alone.
What a pipeline actually ingests
Atlas Bio runs ten live pipelines on fixed cadences, including ClinicalTrials.gov for enrolment status, endpoint changes and results posting; FDA sources for approval actions, labelling changes and post-market adverse event signals; PubMed for publication ingestion into evidence synthesis; and regulatory filings for sponsor pipeline disclosures. Each runs as a scheduled job feeding a calibration module.
The design principle is cadence discipline. A weekly literature sweep and a daily regulatory sweep answer different questions, and a pipeline that quietly stops updating is worse than none at all because the model keeps reporting confidence from stale inputs.
- Trial registry: enrolment, endpoint amendments, results posting
- Regulatory: approvals, labelling changes, post-market safety signals
- Literature: new publications feeding evidence synthesis
- Clinical data partnerships for cohort-level calibration
The limits of multipliers
Atlas Bio publishes three limits on its own approach, and they generalise. Calibration multipliers are population-level, so they cannot identify which patients drive the divergence. They risk double-counting when patient-level inputs such as age or performance status already capture much of the same fragility. And they are static estimates calibrated against historical data, so they must themselves be recalibrated as community practice changes.
The published guidance is to present the trial view and the real-world view side by side rather than relying on either alone, on the basis that the honest answer for any individual patient usually sits between them.
Starting one without a data acquisition budget
Most of the useful early signal is public. Registries, regulatory action databases, labelling histories and the published literature are free and already structured enough to automate, and they cover the questions a small biotech asks first: is anyone else running this design, has the safety language on the comparator changed, what has been published on this mechanism since the last review.
Licensed clinical cohorts become worth their cost later, when the question moves from direction to magnitude — when you need to know not whether real-world outcomes attenuate, but by how much in the specific population you intend to treat.
Atlas Bio runs ten live real-world evidence pipelines across trial registries, regulatory sources and the published literature, feeding calibration multipliers that adjust predictions from trial populations toward community populations.
Is real-world evidence accepted by regulators?
FDA maintains a real-world evidence program established under the 21st Century Cures Act to evaluate the potential use of RWE to support new indications and post-approval study requirements. Acceptance depends on the data's fitness for the specific question and on the analysis methodology.
How much do real-world outcomes differ from trial outcomes?
The direction is consistent — more toxicity, less effectiveness — because trial samples are more selected and lower-risk. The magnitude is indication-specific and should be estimated from comparable real-world cohorts, not assumed from a generic multiplier.
Can adverse event reporting databases give incidence rates?
No. Spontaneous reporting supports post-marketing surveillance and disproportionality analysis, but has no reliable denominator, so it generates hypotheses about signals rather than measuring how often an event occurs.
- FDA. Real-World Evidence — definitions of real-world data and real-world evidence
- FDA. Framework for FDA's Real-World Evidence Program, December 2018
- Kennedy-Martin T et al. A literature review on the representativeness of randomized controlled trial samples and implications for the external validity of trial results. Trials 2015 (PMC4632358)
Talk to Atlas Bio.
Sponsors, researchers and clinicians: tell us the decision you are trying to make and we will show what the platform can and cannot answer.
Contact Us →