Eliza Ward: Brian, I have a confession — I've been low-key annoyed at a statistics textbook for three days and I think I finally figured out why.
Brian Reed: That sounds like our episode right there. What is it?
Eliza Ward: Randomized controlled trials. The whole pitch is: you allocate participants via a coin toss or random number generator — no researcher, no patient can influence the groups — and that breaks confounding. Every lurking variable gets distributed by chance. Clean causal inference. The FDA, NIH, clinical guidelines, all of it, that's the foundation.
Brian Reed: Right, the gold standard. So what's the annoyance?
Eliza Ward: The guarantee is in expectation. Meaning — across many trials, the groups balance. But any single finite-sample trial? Baseline imbalance by pure chance is entirely possible. And we don't run the trial many times. We run it once. The FDA sees one trial. A doctor reads one trial.
Brian Reed: Wait, so the mechanism that justifies the confidence only works in aggregate, and the thing we're actually using is never the aggregate.
Eliza Ward: Yes. And — wait, this is the part that adds historical weight — before the mid-20th century, allocation was based on clinical judgment. Which is an obvious confound: a doctor's preference shapes who gets treated. The RCT revolution fixed that. Genuinely fixed it. But what we built on top is a gold standard claim that's stronger than the underlying probabilistic argument.
Brian Reed: So the question is whether we replaced one kind of certainty problem — bias from judgment — with a quieter one: a one-trial result that the math never promised us we could trust alone.
Eliza Ward: Which is exactly where the coffee thing comes in — because I think that's the click that makes it land.
Brian Reed: Walk me through it.
Eliza Ward: Say you want to know if a new coffee brand makes people more alert. But the people who tend to buy it also tend to sleep better. So the alertness you're seeing — that's not the coffee. That's the sleep. A pre-existing difference that looks like a treatment effect.
Brian Reed: Right, and that's the confound — sleep is associated with both which coffee you pick and how alert you feel, so it creates this spurious apparent causal link that has nothing to do with the coffee itself.
Eliza Ward: Exactly. Randomization is the move that breaks it. Flip a coin to assign the coffee brands — now good sleepers end up roughly equally in both groups. Any alertness difference you see after that? Has to be the coffee. But — and this is the part I kept sliding past — the coin doesn't know what to balance. It doesn't know sleep matters. It just... balances everything. Known confounders, unknown ones, ones you never thought to measure.
Brian Reed: That's the thing that actually got me. It severs all the backdoor pathways from patient characteristics to treatment assignment — not just the ones on your checklist.
Eliza Ward: And there's a practical companion to this that I don't think gets explained enough — allocation concealment. The randomization sequence has to stay hidden until after someone is already enrolled. Because if a researcher can see what the next assignment is, they can subtly steer a healthier patient into the treatment arm. Consciously or not.
Brian Reed: So allocation concealment is — wait, let me get this right — it's not the same as blinding. It's specifically about keeping the sequence hidden so the enrollment decision happens before anyone knows the group. The coin has already landed, you just can't see it yet.
Eliza Ward: Right — and the concealment closes that loophole. But statisticians will tell you that running significance tests on baseline characteristics after randomization is technically inappropriate. Like, formally considered unhelpful. And yet open any published RCT and there's a table — p-values for age, sex, comorbidities, all of it.
Brian Reed: Wait — so the people running the trial know a single random draw might be unbalanced, they're checking for it, and the field says that check is meaningless?
Eliza Ward: That's exactly it. The mechanism only works in aggregate. One trial, one random draw — you might just get unlucky. More older patients in the treatment arm, say. And a p-value on that baseline table doesn't tell you whether that imbalance mattered for your outcome.
Brian Reed: Let me see if I can make that concrete. There's a 68-year-old woman, atrial fibrillation, on warfarin. Her cardiologist has a clean RCT — beautiful internal validity, the observed difference is almost certainly caused by the drug. New anticoagulant, reduced stroke risk. But the trial excluded patients with her level of renal impairment.
Eliza Ward: She wasn't in the room.
Brian Reed: She wasn't in the room. So internal validity — pristine. External validity, for her specifically? I mean, what does it actually reach? Near zero?
Eliza Ward: And — wait, this is the part that surprised me — Peter Rothwell showed in his PLOS Clinical Trials paper that this isn't incidental bad luck. The factors that limit external validity are structural. They're baked into how trials get designed: strict eligibility criteria, controlled environments, homogeneous populations. The gap between internal and external validity isn't a bug in a particular trial. It's a feature of the methodology.
Brian Reed: So the very thing that maximizes confidence the drug caused the effect — narrow population, tight controls — is what makes the result hard to apply to actual patients.
Eliza Ward: And that tension is going to get a lot sharper when we get into what the evidence hierarchy was actually built to answer — because that's where the institutional picture starts to crack.
Brian Reed: And that institutional picture — I mean, the hierarchy didn't just emerge organically. The FDA and NIH built regulatory approval around RCTs. The National Academies of Sciences, Engineering, and Medicine put it in writing — 'Clinical Practice Guidelines We Can Trust' — literally codified RCTs above observational studies in the evidence grading framework.
Eliza Ward: CONSORT is where it gets structural.
Brian Reed: Right — CONSORT means journals won't publish an RCT without specific structural requirements. So the hierarchy isn't just a preference. It's enforced at the infrastructure level. Publication, funding, approval — all of it bends toward the RCT.
Eliza Ward: And I'll defend why that happened — no observational technique fully replicates what randomization does with unmeasured confounders. Not propensity scoring, not instrumental variables. That's not a small caveat. That's the actual reason the hierarchy exists.
Brian Reed: So then what's the problem?
Eliza Ward: The problem is — wait, actually — the hierarchy was built to answer efficacy questions. Controlled conditions, narrow populations. But surgical interventions? You can't blind a surgeon. Rare diseases don't have enough patients to randomize. Long-latency harms take decades. The New England Journal of Medicine ran a major review arguing you need evidence sources beyond RCTs precisely because those categories exist.
Brian Reed: So if surgery, rare diseases, long-term safety — if those collectively represent the majority of real clinical questions—
Eliza Ward: Then the gold standard might apply to a surprisingly narrow slice of what medicine actually does.
Brian Reed: That's the hard fact. The most trusted tool in the hierarchy is structurally inapplicable to whole categories of the questions it's supposed to answer.
Eliza Ward: The hierarchy isn't wrong. It's just — it was built to answer a specific kind of question. Efficacy, controlled conditions, can this drug cause the effect we're measuring. And we've been acting like that specific kind of question is all of medicine.
Brian Reed: And the move the field is starting to make — pragmatic RCTs, broader populations, real-world conditions — that's an attempt to close that gap. But most of the trials that actually go to the FDA for approval are still explanatory designs. Narrow criteria, controlled settings. The regulatory infrastructure hasn't caught up.
Eliza Ward: Right — and I'm not sure it can, fully. Because the tighter the trial, the more certain the causal claim. Loosen it to look like a real Tuesday clinic and you start trading internal validity for external. There's no clean version where you get both.
Brian Reed: So maybe the most honest reframe is — RCTs, observational data, causal inference methods — it's a toolkit. Not a fixed hierarchy. And the actual question is what kind of evidence do you need for this patient, this outcome, this decision.
Eliza Ward: Which is a much harder question than 'did we run an RCT.' Three days ago I was annoyed at a textbook for making randomization sound like a solved problem. I think what actually annoyed me was that the solved problem was a different problem than the one a doctor has.
Brian Reed: Yeah. And I'm not sure we've fully reckoned with what that means. That's — I think that's where I'll sit with it.