June Hadley: Roy, long week — I ended up going to a talk at the medical school on Tuesday, one of those lunchtime things, and someone in the audience asked a question I could not stop thinking about afterward.
Roy Halliday: What was the question?
June Hadley: A clinician stood up and said — I'm paraphrasing — 'the trial says this drug works, but my patient is 67, has diabetes and hypertension and is on six other medications, and every single one of those characteristics was an exclusion criterion in the trial.' And then just sat down.
Roy Halliday: And the room went quiet.
June Hadley: Completely. Because that's not a niche edge case — that is the modal patient. That's the person who actually shows up. And it's the thing I want to really sit with today: what does it mean when the gold standard gives you an answer that doesn't fit the question you're actually asking?
Roy Halliday: The language for this is internal versus external validity. The Randomized Controlled Trial maximizes internal validity — you isolate the causal effect by randomizing. But that isolation is exactly what breaks external validity. The cleaner your controlled trial, the less it looks like the world.
June Hadley: So they're actually in tension — those two things.
Roy Halliday: Directly. And the FDA requires RCT evidence for drug approval. Which means we anchor the entire evidence pyramid to the design that is most likely to produce results that don't generalize.
June Hadley: I mean — that's the thing sitting in that exam room, right? The doctor has a result that is, technically, as rigorous as medicine gets. And it might tell her almost nothing about this specific woman in front of her.
Roy Halliday: A tightly controlled Randomized Controlled Trial can prove causation in a highly selected population while failing entirely to generalize to patients at home with comorbidities and polypharmacy. That's not a flaw in the trial. That's a structural feature of the design.
June Hadley: Which is a distinction I think we mostly just skip over.
Roy Halliday: But here's what that framing skips — the reason the Randomized Controlled Trial became the design in the first place. And it's not arbitrary.
June Hadley: The confounding problem.
Roy Halliday: Exactly that. Before randomization, every study comparing people who exercise to people who don't — contaminated. Because people who exercise also sleep better, earn more, smoke less. The confounding variable is associated with both the exposure and the outcome, but it's not on the causal pathway. You can't see it. You can't measure it if you don't know it's there.
June Hadley: Wait — so it's not just variables you forgot to measure. It's variables you didn't know to look for.
Roy Halliday: That's the core logical move. Randomization doesn't fix the confounders you've identified. It equalizes the ones you haven't. Both arms get the same unknown variables distributed by chance. That's what gives the RCT its internal validity — not measurement, not statistical adjustment. The randomization itself.
June Hadley: So it's not that you measured everything — it's that you didn't need to?
Roy Halliday: Precisely. That's why the FDA's position isn't bureaucratic stubbornness. The evidence pyramid puts systematic reviews of RCTs at the top because no other design can make that move. Observational work — I mean, it can be careful, it can be rigorous — but it cannot randomly assign. So it cannot neutralize the unknown.
June Hadley: Okay, but — does the protection hold through the whole trial? Like, you randomize at the start and then...
Roy Halliday: No, and this is the part people miss. You randomize at assignment. But then some participants don't comply — they switch arms, they drop out. If you analyze only those who actually followed the protocol, that's per-protocol analysis, and you've reintroduced confounding. Because now you're selecting on a characteristic correlated with the outcome. The intention-to-treat analysis exists specifically to preserve the randomization's protection analytically, not just at the design stage.
June Hadley: So the design can be right and the analysis can still break it.
Roy Halliday: Right. And there's a separate problem layered on top — selection bias. That's distinct from confounding. Confounding is about unmeasured variables distorting the effect estimate inside your study. Selection bias is when the people who end up enrolled differ from the population you're trying to say something about. That's the external validity problem. Same trial, two different failure modes.
June Hadley: Which is the woman in the exam room. The trial didn't fail internally — it was just never enrolled to reach her.
Roy Halliday: And that's the clean version of the problem. But there's a harder version underneath it — what about the questions where you can't run the trial at all? Not won't. Can't.
June Hadley: That's actually where I want to go. Because I think — mm, the medical framing makes this feel like a fixable logistics problem. Like 'enroll harder, enroll wider.' But some questions are structurally closed to randomization.
Roy Halliday: Nobody can randomize people to poverty. Nobody can randomize people to smoking. The ethics collapse immediately.
June Hadley: And then — okay, take it further. Think about what LIGO, Virgo, and IceCube were doing during the third observing run, O3. Researchers combining sub-threshold gravitational wave signals with high-energy neutrino signals across three separate instruments. Zero manipulation. They cannot nudge a neutron star merger to see what happens. And yet we're building causal understanding of what drives those events.
Roy Halliday: Nobody thinks astrophysics isn't science.
June Hadley: Right — but the evidence pyramid that puts Randomized Controlled Trials at the top was built inside medicine and then kind of... exported everywhere. And in astrophysics, macroeconomics, evolutionary biology — the RCT isn't lower on the hierarchy. It's not on the hierarchy at all. It's not a tool that exists for those questions.
Roy Halliday: So the pyramid isn't a universal epistemological ranking. It's a domain-specific institutional artifact.
June Hadley: That's the category error. The uncomfortable part for the RCT side — even inside medicine, the internal validity guarantee is conditional. There's an assumption baked into every RCT called SUTVA — Stable Unit Treatment Value Assumption. It requires that one participant's treatment doesn't affect another's outcome. No interference between participants.
Roy Halliday: Which collapses instantly in infectious disease.
June Hadley: Instantly. If I vaccinate you, your neighbors' infection risk drops. That's the whole mechanism. SUTVA is violated by design. And the same thing happens in any networked social intervention — the treatment is spreading through relationships, not sitting cleanly inside individual arms.
Roy Halliday: So the hierarchy isn't universal — it's domain-parochial. And the internal validity guarantee isn't even guaranteed. It's contingent on assumptions that regularly don't hold.
June Hadley: Which makes what Austin Bradford Hill did with the smoking-cancer link — forty years of observational convergence, no trial, couldn't ethically run one — either the strongest possible argument that observational reasoning works, or the most damning evidence of how costly it is to work without randomization. That tension is worth sitting with before we're done.
Roy Halliday: Both readings are honest. Let me give you the story and then you tell me which one survives. Bradford Hill, 1950s. He's working with Richard Doll. No randomization possible — you cannot enroll people into a smoking arm. So he builds criteria: strength of association, consistency across studies, specificity, temporality, biological gradient. The dose-response gradient alone — more cigarettes, more cancer, linearly — that's not what confounders usually produce. Confounders produce flat associations, not gradients.
June Hadley: And temporality — the exposure had to precede the disease.
Roy Halliday: Right. Reverse causality was a live concern — maybe sick people reach for cigarettes. Bradford Hill had to rule that out explicitly because the design gave him no protection against it. That's what an RCT buys you by construction. He had to earn it argumentatively, criterion by criterion.
June Hadley: So the criteria are a formalization of the reasoning you'd need when design-based protection is off the table.
Roy Halliday: Exactly that. And it took roughly forty years of convergent evidence across multiple study types before it moved policy. A twelve-week RCT — if it had been ethical to run — could theoretically have settled it in under two years.
June Hadley: Forty years versus two years. I mean — how many people died in that interval?
Roy Halliday: That's the consequence I want to close with. That gap is not methodological imprecision. It is a body count. So the smoking case is a win for observational reasoning — yes, it worked — and simultaneously a confession about what it costs to work that way.
June Hadley: Is that a win for observational methods, or a confession about their cost? I think — actually, I think it's both, and the fact that we mostly celebrate the win means we've absorbed the wrong lesson.
Roy Halliday: Noted. So what did we build in response?
June Hadley: Judea Pearl. The causal graphical model — the DAG, Directed Acyclic Graph. You draw your assumed causal structure before you touch the data. Arrows represent causal direction. And the graph tells you formally which variables to adjust for, which ones to leave alone, and — this is the part I find striking — it surfaces collider bias, which is a way that adjusting for the wrong variable can actually open a spurious path rather than close one.
Roy Halliday: The do-calculus. Pearl's formalism for what 'intervening on a variable' means mathematically, as distinct from merely observing it. That's the theoretical architecture. The applied revolution is the quasi-experimental methods — propensity score matching, instrumental variables, regression discontinuity, difference-in-differences. All late twentieth, early twenty-first century.
June Hadley: All trying to find variation that nature or policy created, rather than variation the researcher imposed.
Roy Halliday: And every single one of them carries an untestable assumption at its core. The instrumental variable — your instrument has to affect the outcome only through the treatment. Not through any other pathway. You cannot verify that from data. The data cannot tell you whether your instrument is valid. You have to assert it.
June Hadley: So we traded design-based protection for assumption-based protection.
Roy Halliday: And that's a difference in kind, not degree. A well-run RCT's internal validity guarantee doesn't depend on a researcher's judgment call. A well-run instrumental variable study's validity depends entirely on one.
June Hadley: Which puts us back — I think — at the Bradford Hill problem. He was also making judgment calls, criterion by criterion, and it took forty years for the accumulation to be undeniable. The quasi-experimental revolution is faster, more formal. But it's the same structure underneath.
Roy Halliday: The methods improve. The epistemological gap between observational inference and randomization doesn't close. It narrows. The question is whether 'narrowed' is good enough — and that depends entirely on what you're deciding.
June Hadley: That's the room I keep returning to. The exam room. The rheumatologist has the trial on the screen — internally valid, cleanly run — and the 67-year-old is sitting across from her with diabetes and hypertension and six other medications, and every one of those was an exclusion criterion. The trial answered its question. It just wasn't her question.
Roy Halliday: And neither host of this conversation has resolved that. I want to be plain about that. Strong internal validity and weak external validity aren't a problem you fix by better methods. That tradeoff is a permanent structural feature of empirical knowledge. You cannot have both, fully, at once.
June Hadley: Permanent. I think — I mean, that's the honest word for it, isn't it. Not a gap we're closing. A feature.
Roy Halliday: What we actually need isn't a better hierarchy. It's institutions — the FDA, grant committees, journal editors — willing to ask 'which question are you trying to answer?' before they rank the evidence. That's the prior move. Everything else follows from it.
June Hadley: And willing to say both things out loud — we have strong causal proof for this narrow population, and weaker causal proof for everyone else — without collapsing one into the other to sound more certain than we are. That's the honest move. I don't know that we've built institutions that reward honesty that specific.
Roy Halliday: We haven't. Frankly, that's a harder problem than the methodology. Thanks for the Tuesday talk.