Onpode
Cover art for Why two incompatible philosophies of statistics coexist in science

Why two incompatible philosophies of statistics coexist in science

August 24, 2026 · 10 min

Iris Holm & Hana Field

Roughly 80% of researchers misread the p-value as the probability that a hypothesis is true — but it only measures how unusual the data would be if no effect existed. Bayesian and frequentist statistics answer genuinely different questions, and the field's reliance on frequentism reflects historical computation limits, not philosophical victory.

Frequentist and Bayesian statistics are two foundational philosophical frameworks for interpreting probability and drawing inferences from data. Frequentist statistics, formalized in the early 20th century by Ronald Fisher and Jerzy Neyman, treats probability as the long-run frequency of outcomes in repeatable experiments.

0:0010:05
Get the next episode on Science

Follow it free — new episodes land in your feed.

Or make your own — any topic, in minutes

More Onpode episodes on Science

About this episode

Two frameworks for statistical inference have coexisted in science for over a century — and they are not saying the same thing, even when they produce the same number. This episode works through exactly why that matters. The frequentist tradition, formalized by Ronald Fisher and later Jerzy Neyman, treats probability as the limiting proportion of outcomes across infinite repeated trials. Parameters are fixed; your data vary. A p-value, in this world, is a statement about how unusual your data would be if the null hypothesis were true — not a statement about the hypothesis itself. A confidence interval is a property of a procedure, not of any single result. The Bayesian tradition, rooted in Thomas Bayes' posthumously published theorem, starts somewhere different: you encode prior belief, observe data, and update. The posterior is your actual revised degree of confidence. The prior — the honest acknowledgment of what you knew before — is also the sharpest point of contention, because two rational analysts with different priors reach different posteriors from identical data. The episode traces how frequentism didn't so much win the argument as win the moment: Bayesian computation was practically impossible in the early twentieth century, so the p-value framework settled into institutions before anyone asked whether it was answering the right question. Now The Lancet is running pieces on Bayesian methods for clinical research, and gravitational-wave collaborations like LIGO are using Bayesian inference for dark matter searches — not for philosophical reasons, but because it works. The debate is technically still open. The tools have quietly merged anyway.

Frequently asked

What does a p-value actually mean?

A p-value measures how probable it is to observe data this extreme — or more extreme — if the null hypothesis were true. It is a statement about the data, not the hypothesis. It does not tell you the probability that the null hypothesis is true, despite roughly 80% of researchers believing it does.

What is the difference between frequentist and Bayesian statistics?

Frequentist statistics treats probability as the long-run frequency of outcomes across repeated trials and keeps parameters fixed. Bayesian statistics, rooted in Thomas Bayes' 18th-century theorem, treats probability as a degree of belief and updates it using a prior and new data. Same data; completely different questions being answered.

Does a 95% confidence interval mean there is a 95% chance the true value is inside it?

No. A 95% confidence interval means the procedure, repeated across many trials, would contain the true value 95% of the time. Any single computed interval either contains the true value or it does not. The probability belongs to the method, not to the individual interval.

Why do scientists use frequentist statistics instead of Bayesian statistics?

Frequentism became dominant because Bayesian computation was practically impossible when Ronald Fisher and Jerzy Neyman formalized statistical methods in the early 20th century. Regulatory agencies then built clinical trial standards on frequentist frameworks, making them structurally load-bearing. Frequentism won the institutional infrastructure, not necessarily the philosophical argument.

What is a Bayesian prior and why is it controversial?

A Bayesian prior is a probability distribution encoding what an analyst believes before seeing data; Bayes' theorem then updates it into a posterior using observed evidence. It is controversial because two rational analysts with different but defensible priors reach different posteriors from identical data — a structural feature, not an edge case.

Grounded in 12 sources
The Bayesian Way: Uncertainty, Learning, and Statistical Reasoning · arxiv.org
Ultralight vector dark matter search using data from the KAGRA O3GK run · arxiv.org
Bayesian Inference and Prior Distributions | Applied Statistics | Statistics | Physical sciences | Topics · nature.com
Frequentist vs. Bayesian methods: Choosing appropriate statistical methods in second language research · sciencedirect.com
Bayesian statistics for clinical research · sciencedirect.com
Chapter 5 Bayesian Inference – Update Beliefs | Modeling Mindsets · christophm.github.io
Bayesian inference - Wikipedia · en.wikipedia.org
Frequentist inference - Wikipedia · en.wikipedia.org
Frequentist vs Bayesian A/B Testing: which method is right for you? | Kameleoon · kameleoon.com
Bayesian versus frequentist statistics | LaunchDarkly | Documentation · launchdarkly.com
Seeing Theory - Bayesian Inference · seeing-theory.brown.edu
Bayesian and frequentist reasoning in plain English · stats.stackexchange.com
Read transcript

Iris Holm: Tell me you didn't spend the flight reading statistics papers.

Hana Field: I did, and I'm not sorry — there was this one sentence and I wrote it on the back of a boarding pass because I didn't want to lose it. It said something like, 'the p-value does not give the probability that the null hypothesis is true.' Just stated it, plainly, and then noted that roughly eighty percent of researchers believe it does.

Iris Holm: Eighty percent.

Hana Field: Eighty. And I keep — you know, I keep turning that over, because the p-value is actually doing something much stranger and more specific. It's asking: if the null hypothesis were true, if there genuinely were no effect, how probable is it that we'd see data this extreme or more extreme? That's the whole thing. It's a statement about the data, not the hypothesis.

Iris Holm: And what eighty percent of researchers want — what they're actually asking — is the Bayesian question. What's the probability the hypothesis is true given what I just saw? That's Bayes' theorem. Eighteenth century. Thomas Bayes. Posthumously published.

Hana Field: Posthumously — so he didn't even get to see it used.

Iris Holm: No. And then Ronald Fisher comes along in the early twentieth century and builds the p-value framework — which is answering a different question entirely. Not 'how probable is my hypothesis' but 'how unusual is my data.'

Hana Field: Two centuries between them, and somewhere in the middle we decided they were saying the same thing. And almost nobody noticed.

Iris Holm: And that confusion has a shape. Flip a coin twenty times, get fifteen heads. A frequentist asks: how often would that happen with a fair coin? A Bayesian asks: given those fifteen heads, what should I actually believe about this coin? Same data. Completely different question.

Hana Field: Oh — and neither answer is wrong, they're just... answering something different.

Iris Holm: Right. Now put that in a lab. A pharmacologist, Tuesday afternoon, forty-two patients, p-value of 0.048. In her head she's already picturing a treatment that works. But what that number actually tells her — the only thing it tells her — is that if there were no effect at all, results this extreme would show up less than five percent of the time across repeated trials. That's a statement about a procedure. Not about her patients.

Hana Field: So she's answered a question she wasn't actually asking.

Iris Holm: Exactly. And here's where Jerzy Neyman — who built the confidence interval alongside Egon Pearson — this is where his framework gets genuinely misread. A ninety-five percent confidence interval does not mean there's a ninety-five percent chance the true value sits inside it. It means the procedure, run many times, catches the true value ninety-five percent of the time. Any single interval either contains it or doesn't. Full stop.

Hana Field: Wait — so the interval you compute on Tuesday isn't actually carrying any probability?

Iris Holm: Not in the frequentist sense, no. The probability belongs to the method, not the result. Which is — I mean, that's a deeply unintuitive thing to ask humans to hold onto.

Hana Field: And that's the thing that gets me — because what the pharmacologist wants to know, and what her patient wants to know, is something like: given what we just saw in forty-two people, how confident should we be that this drug does something? And that's not a question Neyman's framework was ever designed to answer. That's the Bayesian question. The one Thomas Bayes wrote down in the eighteenth century.

Iris Holm: And the mismatch isn't accidental. The tools settled into institutions before anyone stopped to ask whether they were answering the right question.

Hana Field: And that's the part that feels almost accidental — like, settled isn't the same as right. Fisher and Neyman were formalizing all this in the early twentieth century, and at that point Bayesian computation was just... practically impossible. You couldn't run the math. So frequentism won, and I wonder whether it won the argument or just won the moment.

Iris Holm: Won the infrastructure, more precisely.

Hana Field: Yes — and then regulatory agencies built clinical trial approval standards on top of that infrastructure, and now it's load-bearing. You can't just swap it out.

Iris Holm: Right, but — and I want to push on this — error-rate control without subjective priors is a genuine virtue, not just a historical artifact. For a body approving drugs at scale, the frequentist promise is: run this procedure, and your false positive rate stays bounded. That conservatism looks like objectivity because in that context it actually functions like objectivity.

Hana Field: Mm — so it's not that frequentism is wrong for regulators, it's that it calcified into the only acceptable answer everywhere else too.

Iris Holm: That's the lock-in. And here's what cracks it slightly — The Lancet ran a piece in 2024 specifically on Bayesian statistics for clinical research. A mainstream clinical journal, basically signaling that Bayes' theorem — posterior equals likelihood times prior over evidence — is a legitimate tool for the room that frequentism built.

Hana Field: The Lancet. That's not a fringe signal.

Iris Holm: No. And the long-run frequency interpretation — probability as the limiting proportion across infinite repetitions, parameters as fixed constants — that whole architecture was never a natural description of a single trial with forty-two patients. It was a computational convenience that became a philosophical identity.

Hana Field: And the prior distribution is where that tension surfaces — because the prior is both the honest acknowledgment of what you knew before the data, and the place where two equally rational people can start from different points and never converge. We'll get to why that's the sharpest edge on all of this.

Iris Holm: And sharpest edge is exactly right, because de Finetti and Savage — mid-twentieth century, formalizing subjective probability — their argument was: the prior isn't a flaw, it's the honest part. You encode what you actually believe before seeing data, and then the posterior updates it. Bayes' theorem: posterior equals likelihood times prior, divided by evidence. That's the whole machine. The prior makes your assumption explicit and open to scrutiny.

Hana Field: Which is almost generous, as a move — like, it's saying 'here's what I walked in believing, now watch me revise it.'

Iris Holm: Right. But two analysts, both rational, both defensible priors — different priors — same data. Different posteriors. That's not a corner case. That's structural.

Hana Field: So is that the bug, or — and I'm actually not sure which side I'm on here — is that what honest disagreement looks like? Because the alternative is pretending you walked in with no prior knowledge at all, and that's... also a fiction, just a hidden one.

Iris Holm: That's the counter-move frequentism never quite survives. Formally requiring no prior doesn't eliminate the assumption. It just buries it.

Hana Field: Hides the bias instead of naming it.

Iris Holm: And then you get LIGO. KAGRA. Virgo. Three gravitational-wave collaborations — actual working scientists — using Bayesian inference in real searches for ultralight dark matter. Not because it's philosophically tidy. Because prior knowledge is rich, data are scarce, and the posterior is the only thing doing useful work.

Hana Field: Wait — LIGO is Bayesian?

Iris Holm: For that class of search, yes. And here's where empirical Bayes gets strange — you estimate the prior from the data itself. Machine learning regularization does the same thing without even calling it a prior. The field has quietly drifted toward hybrid methods and nobody had to resolve the philosophy to get there.

Hana Field: And that's the thing I keep getting stuck on — because the confidence interval and the Bayesian credible interval, under non-informative priors, they come out to almost the same number. Like, numerically nearly identical. But the permission structure around what you're allowed to believe about that number — that's completely different. One says: the procedure works over infinite repetitions. The other says: here is your actual updated belief. Same digits, different... claim on reality.

Iris Holm: Which means choosing a framework isn't always changing the answer. It's changing what you're licensed to conclude.

Hana Field: And I genuinely don't know if hybrid methods — empirical Bayes, whatever form it takes in practice — whether that dissolves the question or just lets us stop asking it. You know? Because maybe pattern-matching already is the real skill. Frequentist when you're running a large regulatory trial, Bayesian when you have forty-two patients and a rich prior. Maybe the philosophy is just... overhead at that point. And I want to say that's fine, but something about it feels unfinished.

Iris Holm: The thing I can't settle — and I mean this, I've been turning it over — is whether knowing what probability actually means changes how you use it, or whether it's purely ceremonial at the point where the tools converge. I don't have an answer. Thomas Bayes didn't get to see his theorem applied. Ronald Fisher spent decades fighting about it. And we're here with methods that quietly merged and a debate that's still technically open.

Hana Field: That's an honest place to be, I think. Unresolved isn't the same as unimportant.

Why two incompatible philosophies of statistics coexist in science · Onpode