Onpode
Cover art for Why correlation studies can mislead but are sometimes the only option science has

Why correlation studies can mislead but are sometimes the only option science has

August 2, 2026 · 14 min

Roy Halliday & June Hadley

Randomized controlled trials guarantee internal validity by neutralizing unknown confounders through randomization, but that same design excludes the complex, comorbid patients who actually need treatment. The smoking-lung cancer link took forty years of observational convergence to reach policy — a gap Bradford Hill closed with criteria, not randomization, at a measurable human cost.

Randomized controlled trials (RCTs) are widely regarded as the "gold standard" for establishing causal inference in empirical research. By randomly assigning participants to treatment and control groups, RCTs break the link between participant characteristics and treatment assignment, equalizing both known and unknown confounders and isolating the effect of an intervention.

0:0013:46
Get the next episode on Science

Follow it free — new episodes land in your feed.

Or make your own — any topic, in minutes

More Onpode episodes on Science

About this episode

A clinician at a medical school talk asked a question that stopped the room: if every characteristic of her actual patient was an exclusion criterion in the landmark trial, what exactly does that trial tell her? That question opens a genuinely uncomfortable exploration of what the randomized controlled trial can and cannot do — and why the honest answer is more structural than most science communication admits. The episode works through internal versus external validity and why they are in direct tension, not separate problems to be solved in sequence. It covers how randomization protects against confounders you never knew to look for — and how analysis choices late in a study can quietly undo that protection. It then moves to the harder cases: questions where randomization is ethically impossible, scientifically impossible, or simply the wrong instrument for the domain. From Austin Bradford Hill building a causal case for smoking and cancer criterion by criterion over four decades, to Judea Pearl's Directed Acyclic Graphs formalizing what 'intervention' means mathematically, to the untestable assumptions sitting at the core of every quasi-experimental method — the episode traces how scientists reason toward causation when the gold standard isn't available. The conclusion isn't that observational methods don't work. It's that they work at a cost we tend to celebrate past rather than sit with. Worth your time if you've ever trusted a study without asking who it was actually studying.

Frequently asked

Why can't randomized controlled trials just replace correlation studies entirely?

Randomized controlled trials cannot replace observational studies because some exposures are ethically impossible to assign — researchers cannot randomize people to poverty, smoking, or neutron star mergers. For those questions, observational and quasi-experimental methods are not a lower-ranked option; they are the only option that exists.

What is the difference between internal validity and external validity in clinical trials?

Internal validity means a trial correctly isolates a causal effect in its enrolled population; external validity means those results generalize to real patients. RCTs maximize internal validity through randomization but often exclude comorbid, elderly, or polypharmacy patients — exactly the people clinicians treat — creating a structural tension between the two.

How did Bradford Hill prove smoking causes cancer without a randomized trial?

Austin Bradford Hill and Richard Doll used convergent observational criteria — including strength of association, temporality, and a linear dose-response gradient between cigarette quantity and cancer rate — to establish causation without randomization. Building that case took roughly forty years; an RCT, had one been ethical, could theoretically have settled it in under two years.

What is SUTVA and why does it matter for vaccine and infectious disease trials?

SUTVA — Stable Unit Treatment Value Assumption — requires that one participant's treatment not affect another participant's outcome. It is violated by design in vaccine trials: vaccinating one person reduces neighbors' infection risk. Any networked or social intervention where treatment spreads through relationships also breaks this core RCT assumption.

Can quasi-experimental methods like instrumental variables close the gap with randomized trials?

Quasi-experimental methods such as instrumental variables, regression discontinuity, and difference-in-differences narrow but do not close the gap with RCTs. Every such method carries an untestable assumption at its core — an instrumental variable's validity cannot be verified from data alone — making their protection assumption-based rather than design-based.

Grounded in 10 sources
For objective causal inference, design trumps analysis · arxiv.org
Quantitative probing: Validating causal models using quantitative domain knowledge · arxiv.org
Ultralight vector dark matter search using data from the KAGRA O3GK run · arxiv.org
Deep Search for Joint Sources of Gravitational Waves and High-Energy Neutrinos with IceCube During the Third Observing Run of LIGO and Virgo · arxiv.org
Causal Inference with Observational Data - | C3PO · c3po.media.mit.edu
New Methods for Causal Inference in Randomized Experiments and Observational Studies · dash.harvard.edu
Heterogeneity in a meta-analysis: randomized controlled trials versus observational studies · doi.org
Hunting Causes and Using Them: Approaches in Philosophy and Economics · doi.org
Remarks on Smoking: Evidence and Ethics · doi.org
Causal inference and effect estimation using observational data | Journal of Epidemiology & Community Health · jech.bmj.com
Read transcript

June Hadley: Roy, long week — I ended up going to a talk at the medical school on Tuesday, one of those lunchtime things, and someone in the audience asked a question I could not stop thinking about afterward.

Roy Halliday: What was the question?

June Hadley: A clinician stood up and said — I'm paraphrasing — 'the trial says this drug works, but my patient is 67, has diabetes and hypertension and is on six other medications, and every single one of those characteristics was an exclusion criterion in the trial.' And then just sat down.

Roy Halliday: And the room went quiet.

June Hadley: Completely. Because that's not a niche edge case — that is the modal patient. That's the person who actually shows up. And it's the thing I want to really sit with today: what does it mean when the gold standard gives you an answer that doesn't fit the question you're actually asking?

Roy Halliday: The language for this is internal versus external validity. The Randomized Controlled Trial maximizes internal validity — you isolate the causal effect by randomizing. But that isolation is exactly what breaks external validity. The cleaner your controlled trial, the less it looks like the world.

June Hadley: So they're actually in tension — those two things.

Roy Halliday: Directly. And the FDA requires RCT evidence for drug approval. Which means we anchor the entire evidence pyramid to the design that is most likely to produce results that don't generalize.

June Hadley: I mean — that's the thing sitting in that exam room, right? The doctor has a result that is, technically, as rigorous as medicine gets. And it might tell her almost nothing about this specific woman in front of her.

Roy Halliday: A tightly controlled Randomized Controlled Trial can prove causation in a highly selected population while failing entirely to generalize to patients at home with comorbidities and polypharmacy. That's not a flaw in the trial. That's a structural feature of the design.

June Hadley: Which is a distinction I think we mostly just skip over.

Roy Halliday: But here's what that framing skips — the reason the Randomized Controlled Trial became the design in the first place. And it's not arbitrary.

June Hadley: The confounding problem.

Roy Halliday: Exactly that. Before randomization, every study comparing people who exercise to people who don't — contaminated. Because people who exercise also sleep better, earn more, smoke less. The confounding variable is associated with both the exposure and the outcome, but it's not on the causal pathway. You can't see it. You can't measure it if you don't know it's there.

June Hadley: Wait — so it's not just variables you forgot to measure. It's variables you didn't know to look for.

Roy Halliday: That's the core logical move. Randomization doesn't fix the confounders you've identified. It equalizes the ones you haven't. Both arms get the same unknown variables distributed by chance. That's what gives the RCT its internal validity — not measurement, not statistical adjustment. The randomization itself.

June Hadley: So it's not that you measured everything — it's that you didn't need to?

Roy Halliday: Precisely. That's why the FDA's position isn't bureaucratic stubbornness. The evidence pyramid puts systematic reviews of RCTs at the top because no other design can make that move. Observational work — I mean, it can be careful, it can be rigorous — but it cannot randomly assign. So it cannot neutralize the unknown.

June Hadley: Okay, but — does the protection hold through the whole trial? Like, you randomize at the start and then...

Roy Halliday: No, and this is the part people miss. You randomize at assignment. But then some participants don't comply — they switch arms, they drop out. If you analyze only those who actually followed the protocol, that's per-protocol analysis, and you've reintroduced confounding. Because now you're selecting on a characteristic correlated with the outcome. The intention-to-treat analysis exists specifically to preserve the randomization's protection analytically, not just at the design stage.

June Hadley: So the design can be right and the analysis can still break it.

Roy Halliday: Right. And there's a separate problem layered on top — selection bias. That's distinct from confounding. Confounding is about unmeasured variables distorting the effect estimate inside your study. Selection bias is when the people who end up enrolled differ from the population you're trying to say something about. That's the external validity problem. Same trial, two different failure modes.

June Hadley: Which is the woman in the exam room. The trial didn't fail internally — it was just never enrolled to reach her.

Roy Halliday: And that's the clean version of the problem. But there's a harder version underneath it — what about the questions where you can't run the trial at all? Not won't. Can't.

June Hadley: That's actually where I want to go. Because I think — mm, the medical framing makes this feel like a fixable logistics problem. Like 'enroll harder, enroll wider.' But some questions are structurally closed to randomization.

Roy Halliday: Nobody can randomize people to poverty. Nobody can randomize people to smoking. The ethics collapse immediately.

June Hadley: And then — okay, take it further. Think about what LIGO, Virgo, and IceCube were doing during the third observing run, O3. Researchers combining sub-threshold gravitational wave signals with high-energy neutrino signals across three separate instruments. Zero manipulation. They cannot nudge a neutron star merger to see what happens. And yet we're building causal understanding of what drives those events.

Roy Halliday: Nobody thinks astrophysics isn't science.

June Hadley: Right — but the evidence pyramid that puts Randomized Controlled Trials at the top was built inside medicine and then kind of... exported everywhere. And in astrophysics, macroeconomics, evolutionary biology — the RCT isn't lower on the hierarchy. It's not on the hierarchy at all. It's not a tool that exists for those questions.

Roy Halliday: So the pyramid isn't a universal epistemological ranking. It's a domain-specific institutional artifact.

June Hadley: That's the category error. The uncomfortable part for the RCT side — even inside medicine, the internal validity guarantee is conditional. There's an assumption baked into every RCT called SUTVA — Stable Unit Treatment Value Assumption. It requires that one participant's treatment doesn't affect another's outcome. No interference between participants.

Roy Halliday: Which collapses instantly in infectious disease.

June Hadley: Instantly. If I vaccinate you, your neighbors' infection risk drops. That's the whole mechanism. SUTVA is violated by design. And the same thing happens in any networked social intervention — the treatment is spreading through relationships, not sitting cleanly inside individual arms.

Roy Halliday: So the hierarchy isn't universal — it's domain-parochial. And the internal validity guarantee isn't even guaranteed. It's contingent on assumptions that regularly don't hold.

June Hadley: Which makes what Austin Bradford Hill did with the smoking-cancer link — forty years of observational convergence, no trial, couldn't ethically run one — either the strongest possible argument that observational reasoning works, or the most damning evidence of how costly it is to work without randomization. That tension is worth sitting with before we're done.

Roy Halliday: Both readings are honest. Let me give you the story and then you tell me which one survives. Bradford Hill, 1950s. He's working with Richard Doll. No randomization possible — you cannot enroll people into a smoking arm. So he builds criteria: strength of association, consistency across studies, specificity, temporality, biological gradient. The dose-response gradient alone — more cigarettes, more cancer, linearly — that's not what confounders usually produce. Confounders produce flat associations, not gradients.

June Hadley: And temporality — the exposure had to precede the disease.

Roy Halliday: Right. Reverse causality was a live concern — maybe sick people reach for cigarettes. Bradford Hill had to rule that out explicitly because the design gave him no protection against it. That's what an RCT buys you by construction. He had to earn it argumentatively, criterion by criterion.

June Hadley: So the criteria are a formalization of the reasoning you'd need when design-based protection is off the table.

Roy Halliday: Exactly that. And it took roughly forty years of convergent evidence across multiple study types before it moved policy. A twelve-week RCT — if it had been ethical to run — could theoretically have settled it in under two years.

June Hadley: Forty years versus two years. I mean — how many people died in that interval?

Roy Halliday: That's the consequence I want to close with. That gap is not methodological imprecision. It is a body count. So the smoking case is a win for observational reasoning — yes, it worked — and simultaneously a confession about what it costs to work that way.

June Hadley: Is that a win for observational methods, or a confession about their cost? I think — actually, I think it's both, and the fact that we mostly celebrate the win means we've absorbed the wrong lesson.

Roy Halliday: Noted. So what did we build in response?

June Hadley: Judea Pearl. The causal graphical model — the DAG, Directed Acyclic Graph. You draw your assumed causal structure before you touch the data. Arrows represent causal direction. And the graph tells you formally which variables to adjust for, which ones to leave alone, and — this is the part I find striking — it surfaces collider bias, which is a way that adjusting for the wrong variable can actually open a spurious path rather than close one.

Roy Halliday: The do-calculus. Pearl's formalism for what 'intervening on a variable' means mathematically, as distinct from merely observing it. That's the theoretical architecture. The applied revolution is the quasi-experimental methods — propensity score matching, instrumental variables, regression discontinuity, difference-in-differences. All late twentieth, early twenty-first century.

June Hadley: All trying to find variation that nature or policy created, rather than variation the researcher imposed.

Roy Halliday: And every single one of them carries an untestable assumption at its core. The instrumental variable — your instrument has to affect the outcome only through the treatment. Not through any other pathway. You cannot verify that from data. The data cannot tell you whether your instrument is valid. You have to assert it.

June Hadley: So we traded design-based protection for assumption-based protection.

Roy Halliday: And that's a difference in kind, not degree. A well-run RCT's internal validity guarantee doesn't depend on a researcher's judgment call. A well-run instrumental variable study's validity depends entirely on one.

June Hadley: Which puts us back — I think — at the Bradford Hill problem. He was also making judgment calls, criterion by criterion, and it took forty years for the accumulation to be undeniable. The quasi-experimental revolution is faster, more formal. But it's the same structure underneath.

Roy Halliday: The methods improve. The epistemological gap between observational inference and randomization doesn't close. It narrows. The question is whether 'narrowed' is good enough — and that depends entirely on what you're deciding.

June Hadley: That's the room I keep returning to. The exam room. The rheumatologist has the trial on the screen — internally valid, cleanly run — and the 67-year-old is sitting across from her with diabetes and hypertension and six other medications, and every one of those was an exclusion criterion. The trial answered its question. It just wasn't her question.

Roy Halliday: And neither host of this conversation has resolved that. I want to be plain about that. Strong internal validity and weak external validity aren't a problem you fix by better methods. That tradeoff is a permanent structural feature of empirical knowledge. You cannot have both, fully, at once.

June Hadley: Permanent. I think — I mean, that's the honest word for it, isn't it. Not a gap we're closing. A feature.

Roy Halliday: What we actually need isn't a better hierarchy. It's institutions — the FDA, grant committees, journal editors — willing to ask 'which question are you trying to answer?' before they rank the evidence. That's the prior move. Everything else follows from it.

June Hadley: And willing to say both things out loud — we have strong causal proof for this narrow population, and weaker causal proof for everyone else — without collapsing one into the other to sound more certain than we are. That's the honest move. I don't know that we've built institutions that reward honesty that specific.

Roy Halliday: We haven't. Frankly, that's a harder problem than the methodology. Thanks for the Tuesday talk.

Why correlation studies can mislead but are sometimes the only option science has · Onpode