Onpode
Cover art for Why randomized trials work — the statistical mechanism that isolates cause

Why randomized trials work — the statistical mechanism that isolates cause

September 7, 2026 · 13 min

Marcus Kline & Ben Okonkwo

Randomized controlled trials isolate cause from correlation by randomly assigning who gets treatment, which distributes all confounders — including unknown ones — equally across groups by chance. Austin Bradford Hill applied this logic to medicine in 1948; Ronald Fisher had proved the statistical principle in agricultural experiments decades earlier. The math licenses the causal claim; the ethics constrain where it can go.

Randomized controlled trials (RCTs) are a methodological approach designed to establish causal relationships between interventions and outcomes by randomly assigning participants to treatment and control groups.

0:0013:06
Get the next episode on Science

Follow it free — new episodes land in your feed.

Or make your own — any topic, in minutes

More Onpode episodes on Science

About this episode

In 1948, Austin Bradford Hill handed a life-or-death allocation decision to chance. Half of a tuberculosis ward received streptomycin; the other half became the control. The sealed envelopes were, as this episode puts it, "the most honest thing available" — a procedure for generating deliberate ignorance, which turned out to be the purest form of causal proof medicine had ever seen. But the episode doesn't stop at celebrating the method. It traces the actual statistical machinery: how randomization neutralizes confounding (including the variables no one thought to ask about), why Jerzy Neyman's potential outcomes framework from 1923 and Donald Rubin's formalization in the 1970s arrived long after the method was already trusted, and what it means that the complete mathematical justification was still being written while the trials were running. Then it turns. Rubin's framework, the same architecture that licenses the RCT's causal claim, also extends to observational studies under specific conditions. The math says randomization is sufficient, not uniquely necessary. The profession mostly acted as if it were necessary anyway. That's an institutional choice, not a mathematical conclusion. And the extensions — cluster randomized trials, hybrid RCTs, instrumental variables — each one is a concession to conditions Fisher's crop fields never faced. The episode ends where it has to: in the domains that matter most, climate, education, large-scale social policy, you cannot randomize. The choice isn't gold standard versus inferior method. It's imperfect evidence versus none.

Frequently asked

How do randomized controlled trials eliminate confounding?

Randomized controlled trials eliminate confounding by assigning participants to treatment or control groups by chance. Random assignment distributes every pre-existing difference — age, health status, behavior, and variables no researcher thought to measure — roughly equally across both groups, so the only systematic remaining difference is the intervention itself.

Who invented the randomized controlled trial and when?

Austin Bradford Hill ran the first formal medical RCT in 1948 for the British Medical Research Council, testing streptomycin against pulmonary tuberculosis using sealed envelopes and random sampling numbers. The underlying statistical logic was developed earlier by Ronald Fisher in agricultural experiments in the early twentieth century.

What is the Neyman-Rubin potential outcomes framework?

The Neyman-Rubin potential outcomes framework is the mathematical foundation explaining why RCTs support causal conclusions. Jerzy Neyman introduced potential outcomes in 1923; Donald Rubin formalized the assignment mechanism as a probabilistic model in the 1970s. It holds that every individual has a potential outcome under each treatment, and randomization licenses comparing them.

Can observational studies ever establish causation the way RCTs do?

Donald Rubin's potential outcomes framework shows randomization is a sufficient condition for causal inference, not the only one. Instrumental variable methods can recover causal estimates from observational data when a variable shifts treatment assignment without directly affecting the outcome — but those assumptions must be argued for rather than built in by experimental design.

What are the main limitations of randomized controlled trials?

RCTs face three core limits: small samples can still produce imbalanced groups by chance; blinding is needed separately to prevent placebo effects from contaminating results; and randomization is often ethically or practically impossible at scale — in education, social policy, and climate — making observational methods not inferior but the only available route to causal evidence.

Grounded in 11 sources
Selection Bias in Hybrid Randomized Controlled Trials using External Controls: A Simulation Study · arxiv.org
Using instrumental variables to disentangle treatment and placebo effects in blinded and unblinded randomized clinical trials influenced by unmeasured confounders · arxiv.org
Leveraging contact network structure in the design of cluster randomized trials · arxiv.org
Exploration and Incentivizing Participation in Randomized Trials · arxiv.org
Rethinking the pros and cons of randomized controlled trials and observational studies in the era of big data and advanced methods: a panel discussion · pmc.ncbi.nlm.nih.gov
How to Distinguish Correlation from Causation in Orthopaedic ... · pmc.ncbi.nlm.nih.gov
Understanding and misunderstanding randomized controlled trials · pmc.ncbi.nlm.nih.gov
Randomized Clinical Trial - an overview | ScienceDirect Topics · sciencedirect.com
Correlation vs. Causation: Confounders, Spurious Patterns, and ... · amazon.com
Randomised controlled trial | Better Evaluation · betterevaluation.org
The limitations of randomised controlled trials | CEPR · cepr.org
Read transcript

Marcus Kline: I've been thinking about a specific kind of violence — the epistemic kind. Where the method demands you withhold something you could give, because the proof requires it.

Ben Okonkwo: Okay, I'm in — where are you?

Marcus Kline: Nineteen forty-eight. British Medical Research Council. Pulmonary tuberculosis — patients who will die without treatment. And Austin Bradford Hill deliberately hands the allocation decision to chance. Sealed envelopes. Random sampling numbers. Half of them get streptomycin; half of them become the control.

Ben Okonkwo: And the reason he does that — the actual methodological reason — is to neutralize confounding. Any pre-existing difference between groups: age, baseline health, behavior, things no one even thought to ask about. Randomization distributes all of it, on average, equally. That's the mechanism.

Marcus Kline: But consider what that requires. You know streptomycin might work. You cannot know until the trial ends. And the trial only means something if the coin flip was real. That's the bind.

Ben Okonkwo: And before that trial — before Hill formalizes this for medicine — the problem was that every attempt to demonstrate causation rather than correlation collapsed on the same question: were the patients who received treatment just already different from the ones who didn't?

Marcus Kline: Which is the confounding problem. And randomization is the answer Fisher had already proved in agricultural settings decades before Hill. Early twentieth century — experimental plots, crop yields — Ronald Fisher argues any departure from random assignment introduces experimenter bias and produces inaccurate interpretation.

Ben Okonkwo: Wait — so the core logic is older than the MRC trial by, what, twenty-plus years?

Marcus Kline: The logic, yes. Fisher builds it for crops. Hill carries it to the ward. And the thing that I cannot get past — the thing this episode is really about — is what that act actually invented. It wasn't a measurement tool. It was a procedure for generating ignorance on purpose. Deliberate, controlled not-knowing. And that turns out to be the purest form of proof we have.

Ben Okonkwo: Okay, but I'd push gently on 'purest' — because the formal account of why it works, the Neyman-Rubin potential outcomes framework, that doesn't arrive until much later. Jerzy Neyman introduces potential outcomes in nineteen twenty-three. Donald Rubin formalizes the assignment mechanism as a probabilistic model in the nineteen-seventies. So there's a long window where the method is working, people trust it, and the complete mathematical justification is still being written.

Marcus Kline: Which raises the question I want to sit with for this whole episode. Fisher and Hill — did they understand the full architecture of what they were building? Or did they solve the immediate problem and leave something unexamined?

Ben Okonkwo: That's exactly it. What the sealed envelope doesn't answer — that's where I want to go.

Marcus Kline: What the sealed envelope doesn't answer — that's the thing. So let's go there. Because the justification for why it works comes much later than the method itself.

Ben Okonkwo: Right — and the cleanest way I can frame the actual machinery is this: imagine a student, call her Priya, already highly motivated. A school introduces a new study-hall policy and wants to know if it works. But Priya would go to any study hall, mandatory or not. If she self-selects in, you can't tell whether her grades improve because of the policy or because she was always going to do better anyway.

Marcus Kline: The motivated students cluster in the treatment group.

Ben Okonkwo: Exactly — you've just recreated every observational study failure in one hallway. Now flip a coin. Priya lands in whichever group chance assigns her. So do students who'd never voluntarily show up. Now motivation is scattered across both groups, roughly equally, and the only remaining systematic difference is the policy itself. That's — I mean, that's the whole thing. One coin flip. That's what random assignment does.

Marcus Kline: And Neyman — nineteen twenty-three — he gives that intuition a formal skeleton. Every student, every patient, has a potential outcome under each possible treatment. What they'd look like if treated. What they'd look like if not. The causal effect is the contrast between those two.

Ben Okonkwo: The problem being you can never observe both for the same person simultaneously. Priya either gets the study hall or she doesn't — you can't run both timelines.

Marcus Kline: Which is where Rubin's piece fits. The assignment mechanism — his term — is a probabilistic model of who gets what treatment. In a properly run RCT, that mechanism is fully known and deliberately controlled. And that's what licenses the causal conclusion. Not the measurement. The control over who gets assigned where.

Ben Okonkwo: Now here's where I want to slow down though — because 'on average' is doing enormous work in that sentence and I don't think we should just pass it. Randomization distributes confounders equally on average. Across infinite repetitions of the trial, yes. But in any single trial — say eighty patients in the MRC tuberculosis ward — randomness can still produce imbalanced groups. By chance. That's not a failure of the method; it's what probability actually means.

Marcus Kline: So the MRC trial — eighty patients — that's a bet. A well-structured bet, but still.

Ben Okonkwo: It is. And there's a second gap — even if the groups are balanced, how do you stop a patient who knows they received streptomycin from behaving differently? Resting more, hoping harder. That expectation itself produces measurable improvement. The placebo effect. Blinding is the patch for that — conceal who got what from participants and researchers both, so you're measuring the drug's actual work, not belief in the drug.

Marcus Kline: So the assignment mechanism handles confounding, and blinding handles the gap between observing an outcome and cleanly attributing it to the intervention.

Ben Okonkwo: Yes — and what I think that reveals is that the elegant formal machinery, Neyman's potential outcomes, Rubin's assignment mechanism, it all rests on one assumption: that you fully know and fully control how treatment was assigned. The moment that assumption starts to crack—

Marcus Kline: The whole causal license cracks with it.

Ben Okonkwo: But that crack — that's actually where Rubin does something no one quite anticipated. Because he takes the potential outcomes framework, the same architecture that licenses the RCT's causal claim, and he extends it to observational studies. In the nineteen-seventies. Which means the formal backbone of the gold standard quietly implies... the gold standard isn't uniquely necessary.

Marcus Kline: Hold on. The same Rubin?

Ben Okonkwo: Same framework, yes. What he formalizes is that randomization is one sufficient condition for knowing the assignment mechanism. Not the only possible one. If you can model how treatment was assigned — if you know or can reconstruct the probability that any given person ended up in treatment versus control — you can, under specific conditions, still support a causal conclusion from observational data.

Marcus Kline: Now that troubles me. Because the RCT's power — the thing that makes Hill's sealed envelopes matter — is that randomization controls for unknown confounders simultaneously. Every variable you didn't think to ask about, age, diet, the ward nurse's shift schedule, all of it, neutralized in one coin flip. No observational method can replicate that property. You can only adjust for what you measured.

Ben Okonkwo: That's the position. And it's not wrong — it's just not the whole picture. Because instrumental variable methods exist precisely to do what you're describing: identify a causal effect even when you can't randomize. Elias Chaibub Neto's work, for instance, develops instrumental variables to disentangle treatment effects from placebo effects in unblinded trials. Trials where blinding has already failed.

Marcus Kline: So the instrument is — what, a variable that affects treatment assignment but has no direct path to the outcome?

Ben Okonkwo: Exactly. It's a wedge. You find something that shifts who gets treated but doesn't independently shift the outcome — and you use that variation to recover the causal estimate. It's not randomization. But it's... okay, here's where I want to be careful, because I don't want to overstate it. It's not the same warranty. It relies on assumptions you have to argue for rather than assumptions you built in by design.

Marcus Kline: Which is the gap. Rubin opens the door. He doesn't say how wide it swings.

Ben Okonkwo: And the profession has mostly acted as if the door is nearly closed. Treated RCTs as uniquely privileged, observational methods as inherently inferior. But the formal framework — the math — doesn't actually settle that. That's an institutional choice, not a mathematical conclusion.

Marcus Kline: The participation incentives piece makes that institutional choice look even shakier to me. A patient in a trial who genuinely believes the treatment is superior — they're ethically entitled to prefer it. But if that preference affects whether they enroll or how they behave once assigned, the randomization is compromised. The gold standard cracks on contact with an actual human being.

Ben Okonkwo: That's the exploration-exploitation problem made ethical. The trial needs uniform random exploration across arms. The patient needs the arm they believe is better. Those two things cannot both be fully satisfied at once.

Marcus Kline: And neither of us is resolving that right now — I want to be clear, I don't think there's a clean answer sitting there waiting.

Ben Okonkwo: No. The math and the institution are giving different answers. Rubin's generalization says: randomization is sufficient, not necessary. The profession says: treat it as necessary anyway. Who's right — the framework or the field that built itself around the framework?

Marcus Kline: And that question gets harder when you look at where the RCT actually came from — Fisher's agricultural plots, cheap, numerous, ethically neutral — and ask what happens when you try to carry those conditions somewhere they were never designed to go. That's the thing I think unravels everything else we've said today.

Ben Okonkwo: And that's the crack that opens everything — because Fisher's plots are cheap, numerous, ethically neutral to assign. You can plow under a field. You cannot plow under a person.

Marcus Kline: Which is what Hill inherited without — and I'm not sure he examined it. The conditions that made randomization feasible in Rothamsted's crop fields just... traveled with the method. Unremarked.

Ben Okonkwo: Right — but the part that doesn't fit is what happens when you try to run a cluster randomized trial in, say, a school network. You can't randomize individual students — policy spreads socially, kids talk. So Guy Harling and colleagues developed connectivity-informed CRT designs that actually map the contact network structure first, then randomize. Which sounds like a solution until you flag the assumption it rests on.

Marcus Kline: Independence between clusters.

Ben Okonkwo: Independence between clusters — which the network mapping is actively showing you doesn't exist. You're using the contact structure to design around the violation of the very assumption the design requires.

Marcus Kline: That's — wait. The tool for solving the problem is built from evidence that the problem can't be solved cleanly?

Ben Okonkwo: Roughly, yes. And then hybrid RCTs try a different patch — integrate external control data, reduce the sample size you need. But they cannot guarantee strict type I error rate control. You've traded one bias risk for another. No clean resolution.

Marcus Kline: So every extension — clusters, hybrids — is a concession. Dressed as evolution.

Ben Okonkwo: And the hardest landing: climate, education, social policy at scale — you cannot randomize. Ethics, cost, the sheer scale of it. Which means the actual choice isn't RCT versus observational study. It's imperfect causal evidence versus no evidence at all.

Marcus Kline: No evidence at all.

Ben Okonkwo: That reframes Rubin's generalization completely — it's not just an academic extension of the framework, it's... actually the only route to causal inference in the domains that matter most.

Marcus Kline: And that's the inversion I can't shake. We built the epistemically purest method, and then applied it only where it fits — a controlled ward, a funded trial, an ethical setting — and called everything else inferior. Even when inferior meant impossible.

Ben Okonkwo: The tighter the trial, the more it proves — and the less it resembles the world where the intervention has to work.

Marcus Kline: Which leaves me with the question I don't think either Fisher or Hill ever fully answered. Did they know they were building a proof mechanism for narrow conditions? Or did they believe — genuinely — that it would scale?

Ben Okonkwo: I don't think they knew. And I'm not sure it would have changed anything if they did. Hill had dying patients and a drug that might work. The sealed envelopes were... the most honest thing available in that ward in 1948. The question of whether tuberculosis wards generalize to climate policy — that's not a question you ask when someone is dying in front of you.

Marcus Kline: And that's the part I can't settle. Rubin's framework — the potential outcomes math, the assignment mechanism — it genuinely opens the door to observational evidence as a legitimate causal route. Not a consolation prize. But the institution built around the sealed envelope hasn't moved. The math and the hierarchy are giving different answers, and I don't know which one to trust.

Ben Okonkwo: Neither do I. Honestly.

Marcus Kline: That's — yeah. I think that's where we actually are.

Ben Okonkwo: The envelopes were real. The ward was real. What we inherited along with the method — that's the part still turning over. I'm glad we didn't try to land it cleanly.

Why randomized trials work — the statistical mechanism that isolates cause · Onpode