Onpode
Cover art for How randomization breaks the causation-correlation problem in medical evidence

How randomization breaks the causation-correlation problem in medical evidence

September 1, 2026 · 14 min

Eleanor Crane & Ben Okonkwo

Randomization in medical trials breaks the causation-correlation problem by distributing unknown confounders equally across groups — a method traced to Austin Bradford Hill's 1948 MRC streptomycin trial and Ronald Fisher's agricultural experiments. But the guarantee only holds when participants comply and the trial population matches the real-world patient.

Randomized controlled trials (RCTs) are widely regarded as the methodological gold standard for establishing causal relationships between interventions and outcomes in medicine. The core mechanism is random assignment: by allocating participants to treatment or control groups via chance, RCTs break the link between confounders — both known and unknown — and the treatment received.

0:0013:38
Get the next episode on Health

Follow it free — new episodes land in your feed.

Or make your own — any topic, in minutes

More Onpode episodes on Health

About this episode

In 1948, a statistician named Austin Bradford Hill walked a mathematical proof from agricultural field trials into a tuberculosis ward and changed how medicine decides what works. The proof: let chance assign treatment, and you automatically balance every confounding variable — including the ones you never thought to measure. That's the engine behind the randomized controlled trial, and it's genuinely elegant. But the episode doesn't stop at the founding myth. It traces the specific conditions under which the guarantee holds — and the places it quietly doesn't. Non-compliance inside a trial re-introduces exactly the confounding randomization was designed to destroy. The population enrolled in a landmark trial may exclude the patient sitting in front of you. And the statistical methods designed to repair those problems are, as the episode puts it, a repair kit that isn't finished. Then there's Judea Pearl's challenge: once you understand why randomization works, you can in principle replicate it without the coin flip — by measuring every confounder. And then the Phillips and van der Laan counter: even if you could do that, the inferential mathematics are provably harder. These aren't the same objection. They're different floors of the same building. What holds all of it together is a question about two things 'gold standard' is doing at once: a statistical claim and a trust claim. And those two things are not the same.

Frequently asked

Why does randomization prove causation when observational studies can't?

Randomization proves causation by distributing both known and unknown confounders equally across treatment groups by chance. This breaks the link between who receives treatment and pre-existing differences that could explain outcomes. Fisher formalized this logic in agricultural experiments before Austin Bradford Hill applied it to the 1948 MRC streptomycin tuberculosis trial.

What was the first randomized controlled trial in medicine?

The first published medical randomized controlled trial was the 1948 MRC streptomycin trial for tuberculosis, designed by Austin Bradford Hill. Streptomycin was scarce, and Hill used random assignment — rather than a doctor's judgment — to allocate the drug, founding the modern evidence hierarchy around RCTs.

What is the potential outcomes framework in clinical trials?

The potential outcomes framework holds that every trial participant has two potential outcomes — one if treated, one if not — but only one is ever observed. Randomization makes the treated and untreated groups statistically exchangeable on average, so the difference in observed outcomes is an unbiased estimate of the true causal effect.

How does patient non-compliance undermine randomized controlled trials?

Patient non-compliance re-introduces confounding inside a randomized trial. When participants stop adhering to their assigned treatment, the comparison is no longer between randomly equivalent groups — sicker, anxious, or less stable patients disproportionately stop. Intention-to-treat analysis partially addresses this but estimates a different quantity: the effect of assignment, not of actual treatment.

Can observational studies ever replace RCTs for causal inference?

Judea Pearl argues randomization is not logically irreplaceable: observing and measuring every confounder without error would replicate the same causal guarantee. However, Phillips and van der Laan (2025), drawing on Robins and Ritov (1997), show that statistical inference remains mathematically harder in observational studies even when all confounders are identified — a limitation Pearl's argument does not resolve.

Grounded in 12 sources
When Can Nonrandomized Studies Support Valid Inference ... · ascpt.onlinelibrary.wiley.com
Internal and External Validity Issues in Case Study Research · cambridge.org
Safe inference outside of randomized trials: Application of the stability-controlled quasi-experiment to the effects of three COVID-19 therapies · doi.org
Commentary on ``Nonparametric identification is not enough, but randomized controlled trials are’’: Statistical considerations for generating reliable evidence across a spectrum of studies that increa · doi.org
CAUSAL ANALYSES OF PALLIATIVE CARE OUTCOMES USING OBSERVATIONAL DATA: A REVIEW OF CURRENT LITERATURE · doi.org
Associations in Medical Research Can Be Misleading: A Clinician's Guide to Causal Inference. · doi.org
Semiparametric Estimation of Relative Causal Effects in Randomized Controlled Trials With Noncompliance · doi.org
Toward Generalizing Inferences From Trials to Target Populations · Issue 6.4, Fall 2024 · hdsr.mitpress.mit.edu
Automated causal inference in application to randomized ... · nature.com
Translating evidence into practice: eligibility criteria fail to eliminate clinically significant differences between real-world and study populations · nature.com
The Range and Scientific Value of Randomized Trials: Part 24 of a Series on Evaluation of Scientific Publications - PMC · pmc.ncbi.nlm.nih.gov
Randomized controlled trials – a matter of design · pmc.ncbi.nlm.nih.gov
Read transcript

Eleanor Crane: Ben, I want to start with a scene — 1948, a tuberculosis ward in Britain. Patients are dying. Streptomycin exists but it's scarce. And Austin Bradford Hill, working with the Medical Research Council, does something that had essentially never been done in clinical medicine: he randomly assigns who gets the drug. Not by a doctor's judgment. By chance.

Ben Okonkwo: And that trial — that specific MRC streptomycin-tuberculosis trial — is credited as the first published medical RCT. The founding moment.

Eleanor Crane: Right. And I keep thinking about what Hill was actually afraid of. Because the reason you randomize — the reason chance is fairer than a doctor's hunch — is confounding. The worry that who gets selected for treatment is already different from who doesn't, in ways that corrupt whatever result you see.

Ben Okonkwo: That's the core of it. Confounding is a third variable simultaneously influencing both treatment assignment and outcomes — so in observational data you cannot tell whether the treatment worked or whether you just gave it to people who were going to do better anyway. Randomization breaks that link. That's the mechanism.

Eleanor Crane: And from that one 1948 trial, medicine built an entire hierarchy. The NIH, the Cochrane Collaboration — RCTs at the top, evidence-based medicine structured around that ranking. 'Gold standard' becomes the shorthand.

Ben Okonkwo: Mm, and 'gold standard' is doing interesting work as a phrase — because gold implies purity, finality. But the actual guarantee is narrower than that. The guarantee is: if participants comply with their assigned treatment, and the trial runs as designed, then you have valid causal inference for that population.

Eleanor Crane: Those two conditions — 'if' and 'if.'

Ben Okonkwo: Exactly — and the non-compliance one is the sharper problem, I think. Because once a patient stops taking their medication inside a randomized trial, you're no longer comparing the groups you randomized. Confounding has re-entered the study from the inside. You have a randomized wrapper around what is now, in part, an observational comparison.

Eleanor Crane: And we still hand that result to a clinician and say — this is the highest level of evidence we have.

Ben Okonkwo: We do. Which is not necessarily wrong — but it should make us want to be very precise about what 'gold' is actually promising. That's the question I don't think the label answers.

Eleanor Crane: And what 'gold' is actually promising — that's exactly the thread I want to pull, because I think the promise is more radical than people realize. The coin flip doesn't just balance the things your intake form measured. It balances the things you forgot to put on the form. Things you didn't know to look for.

Ben Okonkwo: That's — yeah, that's the whole move. Unknown confounders. Randomization handles them without naming them.

Eleanor Crane: And that logic didn't start in a hospital. It started in fields. Literally — crop plots.

Ben Okonkwo: Ronald A. Fisher. He was randomizing agricultural experiments — which plot gets which fertilizer treatment — before anyone applied this to patients. And the insight was the same: if soil quality, or drainage, or microclimate is quietly different plot to plot, chance assignment distributes those differences across your groups. You don't need to measure them. You just need to scramble them.

Eleanor Crane: Scramble them — I like that. So when Hill crossed Fisher's logic into that tuberculosis ward in 1948, he was essentially saying: we can't know every reason one patient might do better than another, so we'll let chance absorb all of it.

Ben Okonkwo: Right — and the ethical argument Hill had to make was actually inseparable from the statistical one. A doctor's hunch about who deserves the drug introduces exactly the systematic difference that breaks causal inference. Chance is fairer precisely because it's blind.

Eleanor Crane: Which is almost a strange inversion. Blind is more trustworthy than considered.

Ben Okonkwo: Hm, well — there's a formal name for what the coin flip is actually doing. The potential outcomes framework. The idea is that every participant has two potential outcomes: what happens if they get the treatment, and what happens if they don't. You only ever observe one. But randomization means the group that got treated and the group that didn't are, on average, exchangeable — so the average difference in outcomes between them is an unbiased estimate of the actual causal effect. That's it. That's the engine.

Eleanor Crane: Wait — so you never actually see what the untreated outcome would have been for any individual patient.

Ben Okonkwo: Never. That's the fundamental problem randomization is solving — or, more precisely, routing around. You can't run the counterfactual on a single person. You approximate it across groups made equivalent by chance.

Eleanor Crane: And that's what Fisher proved was mathematically valid — not just intuitively reasonable, but actually provable — through those agricultural experiments. Field by field. Before a single patient was enrolled in anything.

Ben Okonkwo: Exactly. And then Hill takes that proof and walks it into a clinical crisis — scarce drug, dying patients — and the method holds. Fisher's fields were controlled environments. Hill's ward was not. And the gap between those two settings — that's not a footnote. That's where all the interesting questions live.

Eleanor Crane: And who those questions live for — that's what I think gets lost when we stop at the method. The question isn't just whether the logic held in 1948. It's whether it holds for the next patient.

Ben Okonkwo: But that's exactly where the next patient is the problem — because 'holds for the next patient' depends on whether the trial ran clean. And the dirty secret is: perfect compliance is an assumption. It's almost never actually met.

Eleanor Crane: Wait — how far off is 'almost never'?

Ben Okonkwo: Far enough that it matters structurally. The moment a patient stops taking their assigned treatment — not drops out, just stops adhering — you're no longer comparing randomized groups. You're comparing people who comply with people who don't. And those groups, I mean, they differ in ways you cannot fully measure. Sicker patients stop taking things. Anxious patients stop. People with chaotic lives stop. The coin flip assigned them randomly, but their behavior re-sorts them.

Eleanor Crane: The confounder-treatment link that randomization broke — non-compliance partially re-welds it. From the inside.

Ben Okonkwo: Right. And we have tools — intention-to-treat analysis, semiparametric estimation — that are designed to recover valid causal estimates under non-compliance. But here's what I want to be careful about: those methods are still being refined. Semiparametric estimation for this specific problem is, I mean, it's under active development. It's not a settled toolbox.

Eleanor Crane: So the gold standard is — actually — relying on a repair kit that isn't finished.

Ben Okonkwo: That's a harder way to say it than I would. But it's not wrong.

Eleanor Crane: And this is what I keep wanting to press on — because in regulatory submissions, a protocol deviation can disqualify a result. There's a line where the trial is no longer what it claimed to be. But non-compliance, which is structurally doing the same damage to internal validity, gets absorbed into the analysis rather than treated as a failure of the design. Why is that line drawn so differently?

Ben Okonkwo: Hm — I don't think I have a clean answer to that. I mean, intention-to-treat is partly a regulatory compromise: you analyze everyone you randomized, regardless of what they actually did, which protects against certain biases but it also means your estimate is for 'what happens if we assign this treatment in the real world' rather than 'what happens if people actually take it.' Those are different quantities.

Eleanor Crane: Different questions dressed in the same answer.

Ben Okonkwo: Which brings me to a scenario that I think makes this tactile. A 62-year-old woman — diabetes, borderline kidney disease. Her cardiologist recommends a blood pressure drug based on a landmark trial. But that trial excluded patients with her kidney profile. So internal validity inside the trial was real — the design held for those participants. But her doctor is extrapolating to someone the trial was never built to answer for. The internal validity guarantee is already strained before you even ask whether she'll comply.

Eleanor Crane: And she doesn't know that. She's sitting in the office and she's been told this is the proven treatment. And 'proven' is — well, it's conditional on a population that doesn't include her.

Ben Okonkwo: And the deeper problem — which, honestly, we'll have to sit with because it gets harder — is that even the question of whether randomization is the logically irreplaceable part of all this is now contested. Pearl's challenge, the Phillips and van der Laan counter — that's a different layer of the same crack we've just been tracing.

Eleanor Crane: That crack — Pearl's challenge — I want to stay right there. Because I think what Pearl is actually saying is that randomization isn't magic. It's logic. And once you understand the logic, you can, in principle, replicate the guarantee another way.

Ben Okonkwo: His exact words are worth saying out loud: 'once we have understood why RCTs work, there is no need to put them on a pedestal and treat them as the gold standard of causal analysis.' That's not a soft hedge. That's a direct claim.

Eleanor Crane: Which is a remarkable thing for a statistician to say.

Ben Okonkwo: Right — and the logic underneath it is that if you observe every confounder, measure them without error, you've replicated what randomization does mechanically. You've broken the same link. So the RCT isn't irreplaceable; it's just the most reliable way to guarantee that.

Eleanor Crane: Now, the paper that stopped me — Phillips and van der Laan, 2025, drawing on Robins and Ritov from 1997 — they don't dispute the logic. They say the logic is true and practically unreachable.

Ben Okonkwo: Wait — say that more carefully. What do they actually show?

Eleanor Crane: Statistical inference — not just estimation, the inference part — is fundamentally harder in observational studies even when you've observed every confounder. That's not a practical complaint about missing data. It's a mathematical result. The uncertainty doesn't collapse the way it does under randomization.

Ben Okonkwo: Hm. Okay, so — I mean, that's actually a sharper rebuttal than I expected. Because Pearl's argument is about identification. Whether you can, in principle, recover the causal quantity. But Phillips and van der Laan are saying: even granted identification, the inferential problem is harder. Those are different layers.

Eleanor Crane: Two different floors of the same building.

Ben Okonkwo: And then there's the epistemic layer on top of both of those — which is, I think, actually the load-bearing one for clinical practice. You cannot verify, in any specific real-world observational study, that you have actually found all the confounders. Not in principle. In practice, for this patient, this dataset, this moment.

Eleanor Crane: So Pearl's guarantee is — logically valid, empirically unconfirmable.

Ben Okonkwo: Which is why the quasi-experimental methods matter — difference-in-differences, instrumental variables, interrupted time series. They were applied to remdesivir, hydroxychloroquine, dexamethasone in 2025 precisely because you couldn't randomize fast enough or broadly enough. But those methods are still carrying the same epistemic burden. You're arguing your instrument is valid; you can't prove it the way a coin flip proves it.

Eleanor Crane: And that's the thing — I think 'gold standard' is doing two jobs simultaneously and nobody says so. There's a statistical claim: randomization makes inference tractable. And there's a trust claim: when a doctor tells a patient 'this is proven,' they're not citing a p-value. They're invoking an institution.

Ben Okonkwo: And collapsing those two claims is — I mean, that's where the real confusion lives. Because you can weaken the statistical argument with Pearl and the epistemic argument survives. You cannot touch the trust claim with math.

Eleanor Crane: Which means when a patient asks 'is this proven?' — the honest answer is: proven to what standard, for which population, assuming what about compliance. And that's not an answer anyone can give at a bedside.

Ben Okonkwo: No. And that gap — between what the method actually guarantees and what the label 'gold standard' communicates — that's not a technical problem you can design away. It's structural.

Eleanor Crane: I keep thinking about that tuberculosis ward. 1948. Hill randomizes because doctors can't trust their own judgment — their clinical intuition is the problem. The method is a form of institutional humility, almost. And now we've — I mean, have we overcorrected? We trust the method now more than we examine its conditions.

Ben Okonkwo: That's the inversion that's hard to sit with. Hill's argument was: distrust yourself, trust the coin flip. But the coin flip only holds when compliance holds, when the population holds. And neither of those is the coin's job.

Eleanor Crane: Right — but the part that doesn't fit is what a clinician is supposed to do with that. Because the honest answer to 'is this proven?' is now something like: proven internally, for that cohort, under those conditions, assuming adherence. That's a conditional sentence. And conditional sentences are genuinely hard to deliver in an exam room.

Ben Okonkwo: And the Phillips and van der Laan result makes it harder, not easier — because even the alternative path Pearl gestures toward, the observational one, carries inferential costs that don't vanish just because you've measured everything. The uncertainty doesn't resolve. It just moves.

Eleanor Crane: Patients dying, drug scarce, Hill reaches for chance because nothing else was honest enough. That problem hasn't changed. The honesty problem.

Ben Okonkwo: No. It really hasn't. I'm glad we didn't try to wrap it.