Clara Bennett: Hey — before we get going, quick question: if I told you something had been tested thousands of times, failed every single time, and scientists still argue about whether it counts as science, what would you assume I was describing?
Max Rivera: Hmm, I mean — cold fusion? Something in physics that keeps almost working?
Clara Bennett: Homeopathy. Thousands of placebo-controlled trials. Consistent outcome: no effect beyond placebo. And the scientific community calls it pseudoscience — not 'falsified,' pseudoscience.
Max Rivera: Wait — those are meaningfully different things to call it.
Clara Bennett: That distinction is the whole episode. Because Karl Popper's criterion — falsifiability, his answer to what he called the demarcation problem, how you separate science from non-science — says a hypothesis dies when the evidence contradicts it. Homeopathy's 'like cures like' and extreme-dilution doctrines were tested. The evidence contradicted them. So by the criterion's own logic, the science worked.
Max Rivera: But we're not treating it like a hypothesis that lost. We're treating it like a hypothesis that was never really in the running.
Clara Bennett: Right — and astrology sharpens it further. Astrology generates predictions checkable against the actual night sky. It is systematic. It should, on a naive reading of falsifiability, qualify as testable. And it's the paradigm case of pseudoscience.
Max Rivera: So the criterion that Popper built to draw the line... might not be drawing it where we think.
Clara Bennett: Or it's drawing a line, but a different one than advertised. In practice what we want to know is: what is the criterion actually doing, versus what we've been told it does?
Max Rivera: Yeah, and the homeopathy case is weird because — I mean, the falsification happened. Thousands of trials is not nothing. And nothing changed. That should break the model.
Clara Bennett: And that's the exact place Popper started — not with the failure to falsify, but with what makes a claim falsifiable in the first place. The core idea is actually simple: a scientific claim has to specify, in advance, what would prove it wrong.
Max Rivera: Wait, let me just put this in the most boring possible terms, because I think that's actually where it clicks.
Max Rivera: Imagine a friend says: 'I'll believe this diet works if I lose five pounds in a month — and if I don't, I'm done.' That's a bet. There's a number, there's a deadline, there's a condition where she admits she's wrong. Now a different friend says the same diet is amazing, but whenever you ask how you'd know it wasn't working, she says — 'well, your body just needs more time to adjust.' That second sentence? No amount of evidence ever lands. She's made the claim untouchable. That's — I mean, that's the whole thing Popper was pointing at.
Clara Bennett: Exactly that. And the part people miss is where Popper was aiming that idea. He was pushing back against the Vienna Circle — a group of philosophers who thought science meant piling up confirmations. Find enough examples that fit your theory, and the theory is verified. Popper said confirmation is cheap. Anyone can find supporting cases. What's hard — and what's meaningful — is committing ahead of time to what would break the theory.
Max Rivera: Confirmation is cheap — that's a good way to put it.
Clara Bennett: Now, Popper also named the escape hatch that bad theories use. He called them ad hoc hypotheses — modifications made not to generate new predictions, but purely to dodge a result that threatened the theory. Blame the patient's vital force. Say the remedy needs three months. Add a variable. Each move is technically a modification, but none of it produces anything new to test.
Max Rivera: That's the tell. Not that the theory failed — it's that every time it fails, there's a ready-made reason why the failure doesn't count.
Clara Bennett: Right. And in 1934 — when Popper published this — it felt like it had actually solved something. A clean logical criterion. No judgment about content required, just structure: does this claim specify the conditions of its own defeat?
Max Rivera: That's — yeah, I can see why that felt like a breakthrough. You don't have to be an expert in astrology or homeopathy or whatever. You just ask one question about the shape of the claim.
Clara Bennett: One structural question. Which is elegant. But — and this is where it starts to get complicated — the homeopathy case we opened with is, in a way, a falsifiable system that was tested, lost, and kept going. Which means the criterion didn't actually do the enforcement work we'd expect it to.
Max Rivera: So the structure was right — the claim was testable, the test ran, the claim lost — and the thing is still standing. Which means either falsifiability isn't the gate we think it is, or something else is doing the actual work of keeping bad ideas out.
Clara Bennett: And that's the something else. The enforcement work — it's not in the logical structure of the claim. It's in whether the system has a mechanism for actually registering failure. Falsifiable claims plug into a loop: you test, something fails, the theory gets revised or dropped. Unfalsifiable claims route around that loop entirely.
Max Rivera: Wait — so the point isn't the shape of the claim, it's whether failure can actually land somewhere.
Clara Bennett: Exactly. And homeopathy is — it's a perfect case because the loop existed. The trials ran. Failure landed. But the practitioners had a prior system for reinterpreting every negative result before it could count against the doctrine.
Max Rivera: The naturopath scenario gets at this really directly. Patient has chronic knee pain, takes the homeopathic remedy, three months go by — nothing. And the response is: 'your vital force is still clearing blockages, give it another season.' That sentence is — I mean, nothing failed. Formally, right? No failure was registered.
Clara Bennett: The failure was reinterpreted at the point of input. Before it could enter the loop at all.
Max Rivera: Which is — huh. That's different from just moving the goalposts. It's more like the goalposts are made of something that bends around the ball.
Clara Bennett: Now astrology does something structurally similar but the surface looks totally different. Astrology makes genuine, checkable predictions. You can look at a horoscope, look at what happened in someone's week, compare them. The night sky is right there.
Max Rivera: So it's not that astrology makes no predictions.
Clara Bennett: No — it makes plenty. The problem is the interpretive layer wrapped around every prediction. A forecast that comes true confirms the system. One that doesn't? Timing was off, or the chart was misread, or Mercury's position introduced interference. Any result gets accommodated. The community has built — in practice — a structure where no single prediction can ever genuinely refute the doctrine.
Max Rivera: So Popper was actually pointing at the feedback mechanism — but falsifiability as a criterion is kind of a blunt instrument for finding it.
Clara Bennett: That's the honest read. Falsifiability is a proxy for the thing we actually care about: does error-correcting information have anywhere to go? And — this is worth holding onto — what it means is you stop asking 'is this claim falsifiable' and start asking 'does anything in this system incentivize listening to the negative result.'
Max Rivera: And if the answer is no — which, for both homeopathy and astrology, it basically is — then we haven't even gotten to the really uncomfortable version of this problem, which is what happens when the claim is coming from inside physics. That part gets significantly messier.
Clara Bennett: String theory is the place where that gets genuinely uncomfortable. Because here's a case that's not fringe — it's funded, it's in the top journals, it has thousands of working physicists behind it — and it currently makes no predictions testable with any existing technology.
Max Rivera: Wait — none? Like, not 'hard to test.' None.
Clara Bennett: None with current instruments. Same with multiverse cosmology. By a strict Popperian reading, they land in the same category as astrology.
Max Rivera: That's — I mean, that should break something in the model, right? Because physicists aren't treating string theory like astrology. It's not filed under pseudoscience.
Clara Bennett: The defense physicists give is that 'not yet testable because the technology doesn't exist' is different from 'structured to avoid testing.' One is a timing problem. The other is a design problem.
Max Rivera: But — wait, who decides which one it is? That's not a logical distinction. Someone has to make a call.
Clara Bennett: Exactly. That judgment isn't delivered by the criterion. A person makes it. Which is precisely what Larry Laudan argued — falsifiability is neither sufficient on its own nor practically applicable as a bright line. It's a gradient, not a gate.
Max Rivera: And then Imre Lakatos comes in with something that's almost — I mean, it's almost a defense of that messiness. His idea is that scientists legitimately protect the core of a theory while exposing the edges to testing. The 'hard core' stays shielded; the 'protective belt' of auxiliary hypotheses takes the hits.
Clara Bennett: And importantly, Lakatos says that's not bad science. That's how productive research programmes actually work. The problem is — that framing makes it very hard to draw the line between legitimate protection of core commitments and just adding hypotheses to dodge refutation.
Max Rivera: Which is exactly what the homeopathy practitioner is doing with 'vital force clearing blockages.' Technically it's an auxiliary hypothesis.
Clara Bennett: Structurally identical. The difference is social and historical, not logical — which is Kuhn's point entirely. What counts as science shifts with paradigms. It's partly a community decision.
Max Rivera: Seven decades of philosophy and nobody's landed a consensus on where the line actually sits. That's — yeah, that's not a criterion that's doing the heavy lifting we assumed.
Clara Bennett: And then someone tried to automate it.
Max Rivera: Right — Houran and Rodeghier, 2026. They encoded falsifiability into AI systems — neural networks — and got over ninety-five percent accuracy at demarcation detection. Which sounds like a win.
Clara Bennett: Except they called it 'algorithmic pseudo-demarcation.' The system could sort claims, but it couldn't explain why. Black-box output. Epistemologically shallow. We automated the label without automating the reasoning.
Max Rivera: We couldn't even hand this off to a machine. The thing Popper couldn't fully formalize in 1934 — we still can't formalize it now.
Clara Bennett: And the National Academy of Sciences in 1999 — landmark declaration, creation science cannot be meaningfully tested, not science — scholars later pointed out NAS was, in making that call, effectively falsifying creationist claims. The 'unfalsifiable' label was itself contested. Even the clearest institutional application of the criterion turns out to require exactly the kind of judgment the criterion was supposed to replace.
Max Rivera: So falsifiability marks the right goal — it's pointing at the feedback loop, the error-correction — it just can't tell you where the field actually is.
Clara Bennett: When I actually encounter a new medical claim, or a cosmological theory asking for my trust — falsifiability is still the right first question. Can this be proven wrong? That's legitimate. But it doesn't settle anything on its own. Because the next question is immediately: proven wrong by whom? Using what methods? Over what timescale? What counts as a fair result? And those aren't logical questions. Those are questions about credibility. About who I trust to run the test.
Max Rivera: Yeah — Popper handed us the right goal line. He just couldn't hand us the referee.
Clara Bennett: And I don't think that's a failure of the criterion. I've been sitting with that. Maybe it's just — what it actually looks like when a problem is genuinely hard. Popper solved the logical structure in 1934. He didn't claim to solve who gets to be an expert, or what consensus looks like, or how institutions decide when enough negative evidence is enough.
Max Rivera: So we're not left with nothing. We're left with — I mean, a real question. Not a broken tool.
Clara Bennett: A real question that requires real judgment every time. That's probably the honest place.
Max Rivera: This was a good one to sit inside for a while. Thank you for that.