Eleanor Crane: Ben, I want to start with a scene — 1948, a tuberculosis ward in Britain. Patients are dying. Streptomycin exists but it's scarce. And Austin Bradford Hill, working with the Medical Research Council, does something that had essentially never been done in clinical medicine: he randomly assigns who gets the drug. Not by a doctor's judgment. By chance.
Ben Okonkwo: And that trial — that specific MRC streptomycin-tuberculosis trial — is credited as the first published medical RCT. The founding moment.
Eleanor Crane: Right. And I keep thinking about what Hill was actually afraid of. Because the reason you randomize — the reason chance is fairer than a doctor's hunch — is confounding. The worry that who gets selected for treatment is already different from who doesn't, in ways that corrupt whatever result you see.
Ben Okonkwo: That's the core of it. Confounding is a third variable simultaneously influencing both treatment assignment and outcomes — so in observational data you cannot tell whether the treatment worked or whether you just gave it to people who were going to do better anyway. Randomization breaks that link. That's the mechanism.
Eleanor Crane: And from that one 1948 trial, medicine built an entire hierarchy. The NIH, the Cochrane Collaboration — RCTs at the top, evidence-based medicine structured around that ranking. 'Gold standard' becomes the shorthand.
Ben Okonkwo: Mm, and 'gold standard' is doing interesting work as a phrase — because gold implies purity, finality. But the actual guarantee is narrower than that. The guarantee is: if participants comply with their assigned treatment, and the trial runs as designed, then you have valid causal inference for that population.
Eleanor Crane: Those two conditions — 'if' and 'if.'
Ben Okonkwo: Exactly — and the non-compliance one is the sharper problem, I think. Because once a patient stops taking their medication inside a randomized trial, you're no longer comparing the groups you randomized. Confounding has re-entered the study from the inside. You have a randomized wrapper around what is now, in part, an observational comparison.
Eleanor Crane: And we still hand that result to a clinician and say — this is the highest level of evidence we have.
Ben Okonkwo: We do. Which is not necessarily wrong — but it should make us want to be very precise about what 'gold' is actually promising. That's the question I don't think the label answers.
Eleanor Crane: And what 'gold' is actually promising — that's exactly the thread I want to pull, because I think the promise is more radical than people realize. The coin flip doesn't just balance the things your intake form measured. It balances the things you forgot to put on the form. Things you didn't know to look for.
Ben Okonkwo: That's — yeah, that's the whole move. Unknown confounders. Randomization handles them without naming them.
Eleanor Crane: And that logic didn't start in a hospital. It started in fields. Literally — crop plots.
Ben Okonkwo: Ronald A. Fisher. He was randomizing agricultural experiments — which plot gets which fertilizer treatment — before anyone applied this to patients. And the insight was the same: if soil quality, or drainage, or microclimate is quietly different plot to plot, chance assignment distributes those differences across your groups. You don't need to measure them. You just need to scramble them.
Eleanor Crane: Scramble them — I like that. So when Hill crossed Fisher's logic into that tuberculosis ward in 1948, he was essentially saying: we can't know every reason one patient might do better than another, so we'll let chance absorb all of it.
Ben Okonkwo: Right — and the ethical argument Hill had to make was actually inseparable from the statistical one. A doctor's hunch about who deserves the drug introduces exactly the systematic difference that breaks causal inference. Chance is fairer precisely because it's blind.
Eleanor Crane: Which is almost a strange inversion. Blind is more trustworthy than considered.
Ben Okonkwo: Hm, well — there's a formal name for what the coin flip is actually doing. The potential outcomes framework. The idea is that every participant has two potential outcomes: what happens if they get the treatment, and what happens if they don't. You only ever observe one. But randomization means the group that got treated and the group that didn't are, on average, exchangeable — so the average difference in outcomes between them is an unbiased estimate of the actual causal effect. That's it. That's the engine.
Eleanor Crane: Wait — so you never actually see what the untreated outcome would have been for any individual patient.
Ben Okonkwo: Never. That's the fundamental problem randomization is solving — or, more precisely, routing around. You can't run the counterfactual on a single person. You approximate it across groups made equivalent by chance.
Eleanor Crane: And that's what Fisher proved was mathematically valid — not just intuitively reasonable, but actually provable — through those agricultural experiments. Field by field. Before a single patient was enrolled in anything.
Ben Okonkwo: Exactly. And then Hill takes that proof and walks it into a clinical crisis — scarce drug, dying patients — and the method holds. Fisher's fields were controlled environments. Hill's ward was not. And the gap between those two settings — that's not a footnote. That's where all the interesting questions live.
Eleanor Crane: And who those questions live for — that's what I think gets lost when we stop at the method. The question isn't just whether the logic held in 1948. It's whether it holds for the next patient.
Ben Okonkwo: But that's exactly where the next patient is the problem — because 'holds for the next patient' depends on whether the trial ran clean. And the dirty secret is: perfect compliance is an assumption. It's almost never actually met.
Eleanor Crane: Wait — how far off is 'almost never'?
Ben Okonkwo: Far enough that it matters structurally. The moment a patient stops taking their assigned treatment — not drops out, just stops adhering — you're no longer comparing randomized groups. You're comparing people who comply with people who don't. And those groups, I mean, they differ in ways you cannot fully measure. Sicker patients stop taking things. Anxious patients stop. People with chaotic lives stop. The coin flip assigned them randomly, but their behavior re-sorts them.
Eleanor Crane: The confounder-treatment link that randomization broke — non-compliance partially re-welds it. From the inside.
Ben Okonkwo: Right. And we have tools — intention-to-treat analysis, semiparametric estimation — that are designed to recover valid causal estimates under non-compliance. But here's what I want to be careful about: those methods are still being refined. Semiparametric estimation for this specific problem is, I mean, it's under active development. It's not a settled toolbox.
Eleanor Crane: So the gold standard is — actually — relying on a repair kit that isn't finished.
Ben Okonkwo: That's a harder way to say it than I would. But it's not wrong.
Eleanor Crane: And this is what I keep wanting to press on — because in regulatory submissions, a protocol deviation can disqualify a result. There's a line where the trial is no longer what it claimed to be. But non-compliance, which is structurally doing the same damage to internal validity, gets absorbed into the analysis rather than treated as a failure of the design. Why is that line drawn so differently?
Ben Okonkwo: Hm — I don't think I have a clean answer to that. I mean, intention-to-treat is partly a regulatory compromise: you analyze everyone you randomized, regardless of what they actually did, which protects against certain biases but it also means your estimate is for 'what happens if we assign this treatment in the real world' rather than 'what happens if people actually take it.' Those are different quantities.
Eleanor Crane: Different questions dressed in the same answer.
Ben Okonkwo: Which brings me to a scenario that I think makes this tactile. A 62-year-old woman — diabetes, borderline kidney disease. Her cardiologist recommends a blood pressure drug based on a landmark trial. But that trial excluded patients with her kidney profile. So internal validity inside the trial was real — the design held for those participants. But her doctor is extrapolating to someone the trial was never built to answer for. The internal validity guarantee is already strained before you even ask whether she'll comply.
Eleanor Crane: And she doesn't know that. She's sitting in the office and she's been told this is the proven treatment. And 'proven' is — well, it's conditional on a population that doesn't include her.
Ben Okonkwo: And the deeper problem — which, honestly, we'll have to sit with because it gets harder — is that even the question of whether randomization is the logically irreplaceable part of all this is now contested. Pearl's challenge, the Phillips and van der Laan counter — that's a different layer of the same crack we've just been tracing.
Eleanor Crane: That crack — Pearl's challenge — I want to stay right there. Because I think what Pearl is actually saying is that randomization isn't magic. It's logic. And once you understand the logic, you can, in principle, replicate the guarantee another way.
Ben Okonkwo: His exact words are worth saying out loud: 'once we have understood why RCTs work, there is no need to put them on a pedestal and treat them as the gold standard of causal analysis.' That's not a soft hedge. That's a direct claim.
Eleanor Crane: Which is a remarkable thing for a statistician to say.
Ben Okonkwo: Right — and the logic underneath it is that if you observe every confounder, measure them without error, you've replicated what randomization does mechanically. You've broken the same link. So the RCT isn't irreplaceable; it's just the most reliable way to guarantee that.
Eleanor Crane: Now, the paper that stopped me — Phillips and van der Laan, 2025, drawing on Robins and Ritov from 1997 — they don't dispute the logic. They say the logic is true and practically unreachable.
Ben Okonkwo: Wait — say that more carefully. What do they actually show?
Eleanor Crane: Statistical inference — not just estimation, the inference part — is fundamentally harder in observational studies even when you've observed every confounder. That's not a practical complaint about missing data. It's a mathematical result. The uncertainty doesn't collapse the way it does under randomization.
Ben Okonkwo: Hm. Okay, so — I mean, that's actually a sharper rebuttal than I expected. Because Pearl's argument is about identification. Whether you can, in principle, recover the causal quantity. But Phillips and van der Laan are saying: even granted identification, the inferential problem is harder. Those are different layers.
Eleanor Crane: Two different floors of the same building.
Ben Okonkwo: And then there's the epistemic layer on top of both of those — which is, I think, actually the load-bearing one for clinical practice. You cannot verify, in any specific real-world observational study, that you have actually found all the confounders. Not in principle. In practice, for this patient, this dataset, this moment.
Eleanor Crane: So Pearl's guarantee is — logically valid, empirically unconfirmable.
Ben Okonkwo: Which is why the quasi-experimental methods matter — difference-in-differences, instrumental variables, interrupted time series. They were applied to remdesivir, hydroxychloroquine, dexamethasone in 2025 precisely because you couldn't randomize fast enough or broadly enough. But those methods are still carrying the same epistemic burden. You're arguing your instrument is valid; you can't prove it the way a coin flip proves it.
Eleanor Crane: And that's the thing — I think 'gold standard' is doing two jobs simultaneously and nobody says so. There's a statistical claim: randomization makes inference tractable. And there's a trust claim: when a doctor tells a patient 'this is proven,' they're not citing a p-value. They're invoking an institution.
Ben Okonkwo: And collapsing those two claims is — I mean, that's where the real confusion lives. Because you can weaken the statistical argument with Pearl and the epistemic argument survives. You cannot touch the trust claim with math.
Eleanor Crane: Which means when a patient asks 'is this proven?' — the honest answer is: proven to what standard, for which population, assuming what about compliance. And that's not an answer anyone can give at a bedside.
Ben Okonkwo: No. And that gap — between what the method actually guarantees and what the label 'gold standard' communicates — that's not a technical problem you can design away. It's structural.
Eleanor Crane: I keep thinking about that tuberculosis ward. 1948. Hill randomizes because doctors can't trust their own judgment — their clinical intuition is the problem. The method is a form of institutional humility, almost. And now we've — I mean, have we overcorrected? We trust the method now more than we examine its conditions.
Ben Okonkwo: That's the inversion that's hard to sit with. Hill's argument was: distrust yourself, trust the coin flip. But the coin flip only holds when compliance holds, when the population holds. And neither of those is the coin's job.
Eleanor Crane: Right — but the part that doesn't fit is what a clinician is supposed to do with that. Because the honest answer to 'is this proven?' is now something like: proven internally, for that cohort, under those conditions, assuming adherence. That's a conditional sentence. And conditional sentences are genuinely hard to deliver in an exam room.
Ben Okonkwo: And the Phillips and van der Laan result makes it harder, not easier — because even the alternative path Pearl gestures toward, the observational one, carries inferential costs that don't vanish just because you've measured everything. The uncertainty doesn't resolve. It just moves.
Eleanor Crane: Patients dying, drug scarce, Hill reaches for chance because nothing else was honest enough. That problem hasn't changed. The honesty problem.
Ben Okonkwo: No. It really hasn't. I'm glad we didn't try to wrap it.