Iris Holm: Tell me you didn't spend the flight reading statistics papers.
Hana Field: I did, and I'm not sorry — there was this one sentence and I wrote it on the back of a boarding pass because I didn't want to lose it. It said something like, 'the p-value does not give the probability that the null hypothesis is true.' Just stated it, plainly, and then noted that roughly eighty percent of researchers believe it does.
Iris Holm: Eighty percent.
Hana Field: Eighty. And I keep — you know, I keep turning that over, because the p-value is actually doing something much stranger and more specific. It's asking: if the null hypothesis were true, if there genuinely were no effect, how probable is it that we'd see data this extreme or more extreme? That's the whole thing. It's a statement about the data, not the hypothesis.
Iris Holm: And what eighty percent of researchers want — what they're actually asking — is the Bayesian question. What's the probability the hypothesis is true given what I just saw? That's Bayes' theorem. Eighteenth century. Thomas Bayes. Posthumously published.
Hana Field: Posthumously — so he didn't even get to see it used.
Iris Holm: No. And then Ronald Fisher comes along in the early twentieth century and builds the p-value framework — which is answering a different question entirely. Not 'how probable is my hypothesis' but 'how unusual is my data.'
Hana Field: Two centuries between them, and somewhere in the middle we decided they were saying the same thing. And almost nobody noticed.
Iris Holm: And that confusion has a shape. Flip a coin twenty times, get fifteen heads. A frequentist asks: how often would that happen with a fair coin? A Bayesian asks: given those fifteen heads, what should I actually believe about this coin? Same data. Completely different question.
Hana Field: Oh — and neither answer is wrong, they're just... answering something different.
Iris Holm: Right. Now put that in a lab. A pharmacologist, Tuesday afternoon, forty-two patients, p-value of 0.048. In her head she's already picturing a treatment that works. But what that number actually tells her — the only thing it tells her — is that if there were no effect at all, results this extreme would show up less than five percent of the time across repeated trials. That's a statement about a procedure. Not about her patients.
Hana Field: So she's answered a question she wasn't actually asking.
Iris Holm: Exactly. And here's where Jerzy Neyman — who built the confidence interval alongside Egon Pearson — this is where his framework gets genuinely misread. A ninety-five percent confidence interval does not mean there's a ninety-five percent chance the true value sits inside it. It means the procedure, run many times, catches the true value ninety-five percent of the time. Any single interval either contains it or doesn't. Full stop.
Hana Field: Wait — so the interval you compute on Tuesday isn't actually carrying any probability?
Iris Holm: Not in the frequentist sense, no. The probability belongs to the method, not the result. Which is — I mean, that's a deeply unintuitive thing to ask humans to hold onto.
Hana Field: And that's the thing that gets me — because what the pharmacologist wants to know, and what her patient wants to know, is something like: given what we just saw in forty-two people, how confident should we be that this drug does something? And that's not a question Neyman's framework was ever designed to answer. That's the Bayesian question. The one Thomas Bayes wrote down in the eighteenth century.
Iris Holm: And the mismatch isn't accidental. The tools settled into institutions before anyone stopped to ask whether they were answering the right question.
Hana Field: And that's the part that feels almost accidental — like, settled isn't the same as right. Fisher and Neyman were formalizing all this in the early twentieth century, and at that point Bayesian computation was just... practically impossible. You couldn't run the math. So frequentism won, and I wonder whether it won the argument or just won the moment.
Iris Holm: Won the infrastructure, more precisely.
Hana Field: Yes — and then regulatory agencies built clinical trial approval standards on top of that infrastructure, and now it's load-bearing. You can't just swap it out.
Iris Holm: Right, but — and I want to push on this — error-rate control without subjective priors is a genuine virtue, not just a historical artifact. For a body approving drugs at scale, the frequentist promise is: run this procedure, and your false positive rate stays bounded. That conservatism looks like objectivity because in that context it actually functions like objectivity.
Hana Field: Mm — so it's not that frequentism is wrong for regulators, it's that it calcified into the only acceptable answer everywhere else too.
Iris Holm: That's the lock-in. And here's what cracks it slightly — The Lancet ran a piece in 2024 specifically on Bayesian statistics for clinical research. A mainstream clinical journal, basically signaling that Bayes' theorem — posterior equals likelihood times prior over evidence — is a legitimate tool for the room that frequentism built.
Hana Field: The Lancet. That's not a fringe signal.
Iris Holm: No. And the long-run frequency interpretation — probability as the limiting proportion across infinite repetitions, parameters as fixed constants — that whole architecture was never a natural description of a single trial with forty-two patients. It was a computational convenience that became a philosophical identity.
Hana Field: And the prior distribution is where that tension surfaces — because the prior is both the honest acknowledgment of what you knew before the data, and the place where two equally rational people can start from different points and never converge. We'll get to why that's the sharpest edge on all of this.
Iris Holm: And sharpest edge is exactly right, because de Finetti and Savage — mid-twentieth century, formalizing subjective probability — their argument was: the prior isn't a flaw, it's the honest part. You encode what you actually believe before seeing data, and then the posterior updates it. Bayes' theorem: posterior equals likelihood times prior, divided by evidence. That's the whole machine. The prior makes your assumption explicit and open to scrutiny.
Hana Field: Which is almost generous, as a move — like, it's saying 'here's what I walked in believing, now watch me revise it.'
Iris Holm: Right. But two analysts, both rational, both defensible priors — different priors — same data. Different posteriors. That's not a corner case. That's structural.
Hana Field: So is that the bug, or — and I'm actually not sure which side I'm on here — is that what honest disagreement looks like? Because the alternative is pretending you walked in with no prior knowledge at all, and that's... also a fiction, just a hidden one.
Iris Holm: That's the counter-move frequentism never quite survives. Formally requiring no prior doesn't eliminate the assumption. It just buries it.
Hana Field: Hides the bias instead of naming it.
Iris Holm: And then you get LIGO. KAGRA. Virgo. Three gravitational-wave collaborations — actual working scientists — using Bayesian inference in real searches for ultralight dark matter. Not because it's philosophically tidy. Because prior knowledge is rich, data are scarce, and the posterior is the only thing doing useful work.
Hana Field: Wait — LIGO is Bayesian?
Iris Holm: For that class of search, yes. And here's where empirical Bayes gets strange — you estimate the prior from the data itself. Machine learning regularization does the same thing without even calling it a prior. The field has quietly drifted toward hybrid methods and nobody had to resolve the philosophy to get there.
Hana Field: And that's the thing I keep getting stuck on — because the confidence interval and the Bayesian credible interval, under non-informative priors, they come out to almost the same number. Like, numerically nearly identical. But the permission structure around what you're allowed to believe about that number — that's completely different. One says: the procedure works over infinite repetitions. The other says: here is your actual updated belief. Same digits, different... claim on reality.
Iris Holm: Which means choosing a framework isn't always changing the answer. It's changing what you're licensed to conclude.
Hana Field: And I genuinely don't know if hybrid methods — empirical Bayes, whatever form it takes in practice — whether that dissolves the question or just lets us stop asking it. You know? Because maybe pattern-matching already is the real skill. Frequentist when you're running a large regulatory trial, Bayesian when you have forty-two patients and a rich prior. Maybe the philosophy is just... overhead at that point. And I want to say that's fine, but something about it feels unfinished.
Iris Holm: The thing I can't settle — and I mean this, I've been turning it over — is whether knowing what probability actually means changes how you use it, or whether it's purely ceremonial at the point where the tools converge. I don't have an answer. Thomas Bayes didn't get to see his theorem applied. Ronald Fisher spent decades fighting about it. And we're here with methods that quietly merged and a debate that's still technically open.
Hana Field: That's an honest place to be, I think. Unresolved isn't the same as unimportant.