Onpode
Cover art for The structural incentives that make false positives publishable but replications unrewarding

The structural incentives that make false positives publishable but replications unrewarding

August 17, 2026 · 9 min

Eleanor Crane & Ben Okonkwo

The replication crisis in science is structurally manufactured: when the Open Science Collaboration retested 100 published psychology studies, only 36 replicated. Journals reject null results, hiring committees reward novelty, and replication carries no career payoff — making false positives the rational output of an incentive system designed wrong.

The replication crisis — also called the reproducibility crisis — refers to widespread failures to reproduce published scientific findings, particularly in psychology, social sciences, and medicine. The problem gained significant public and academic attention in the early-to-mid 2010s.

0:009:25
Get the next episode on Science

Follow it free — new episodes land in your feed.

Or make your own — any topic, in minutes

More Onpode episodes on Science

About this episode

The replication crisis is usually told as a story about bad actors — researchers who fudged numbers, journals that looked away. This episode tells a different story: what if the system is functioning more or less as intended, and that's the problem? The Open Science Collaboration replicated 100 published psychology studies. Thirty-six held up. The episode works through how that number gets produced — not through fraud, but through the compounding of three ordinary practices: p-hacking, HARKing (hypothesizing after results are known), and the multiple comparisons problem. Each one is defensible in isolation; stacked together, they make a false positive nearly inevitable and nearly invisible. Then there's the file drawer. Null results don't get published, so the literature isn't a record of what's true — it's a biased sample of what was surprising. Ego-depletion looked like consensus until publication-bias-corrected analyses found no reliable effect at all. In cancer biology, landmark pre-clinical findings replicated at 11%. The episode also sits with the attempted fixes — Registered Reports, pre-registration, the Open Science Framework — and asks whether they change behavior or just paperwork, when hiring committees still count Nature papers and journals still chase novelty. No clean answer. Worth the nine minutes to think it through.

Frequently asked

How many psychology studies failed to replicate in the Open Science Collaboration project?

The Open Science Collaboration retested 100 published psychology studies and found that only 36 held up — meaning 64 published findings did not survive a direct replication attempt. The project is the most cited large-scale evidence for the scale of psychology's replication crisis.

What is p-hacking and why does it produce false positives?

P-hacking is the practice of repeatedly adjusting an analysis — shifting variables, peeking at results, changing inclusion criteria — until a statistically significant result appears, then reporting only that version. Because multiple tests inflate false-positive rates mathematically, a significant-looking finding can emerge by chance alone, without any deliberate fraud.

What is HARKing in research and why is it a problem?

HARKing — Hypothesizing After Results are Known — means a researcher runs an analysis, observes what emerged, and then writes the hypothesis as if it had been predicted in advance. The resulting paper reads as a clean confirmed prediction, but the hypothesis was drawn around the result after the fact, making it undetectable from outside the study.

Why don't scientists just run more replication studies to fix the replication crisis?

Replication studies face two structural barriers: journals are biased against publishing null results, and academic hiring committees offer no meaningful career reward for running replications. A researcher can build a career on a novel false positive but loses years publishing a true negative — so the corrective mechanism has no professional engine behind it.

Do registered reports and pre-registration actually fix the replication crisis?

Registered Reports — adopted by journals including PLOS Biology and Royal Society Open Science — peer-review study designs before data collection and provisionally accept papers regardless of outcome, which closes the p-hacking and HARKing loops in principle. Critics note, however, that if hiring committees still reward novelty and Nature papers, pre-registration may change the paperwork without changing the underlying incentives.

Grounded in 12 sources
p-Hacking Inflates Type I Error Rates in the Error Statistical Approach but not in the Formal Inference Approach · arxiv.org
Open Badges: A Low-Cost Toolkit for Measuring Team Communication and Dynamics · arxiv.org
Dead Science Walking: Publication Bias and theAI Scientist Pipeline · arxiv.org
Modelling publication bias and p-hacking · arxiv.org
Type I Error Rates are Not Usually Inflated · arxiv.org
Replicability - Reproducibility and Replicability in Science - NCBI Bookshelf · ncbi.nlm.nih.gov
Concerns About Replicability Across Two Crises in Social Psychology · pmc.ncbi.nlm.nih.gov
Campbell’s Law Explains the Replication Crisis: Pre-Registration Badges Are History Repeating · pmc.ncbi.nlm.nih.gov
Publication bias and the canonization of false facts · pmc.ncbi.nlm.nih.gov
Open Science Practices in Personality Disorder Journals - PMC · pmc.ncbi.nlm.nih.gov
Unlocking the benefits of transparent and reusable science ... · pnas.org
Pre-registration: Weighing costs and benefits for researchers · sciencedirect.com
Read transcript

Eleanor Crane: You know what I noticed reading through this — I never once asked myself whether the scientists were honest. I just kept asking what they were being paid to do.

Ben Okonkwo: That's — yeah, that's the entry point, actually. Because the data doesn't flatter anyone, but it also doesn't accuse anyone. A hundred published psychology studies. The Open Science Collaboration replicates all of them. Thirty-six hold up.

Eleanor Crane: Sixty-four published findings that simply don't survive a second look.

Ben Okonkwo: And I want to flag the assumption the whole crisis narrative rests on — because some researchers, Stroebe and Strack among them, pushed back and said failed replications might just mean we missed a moderator, some contextual variable the replication didn't match. Which is not nothing.

Eleanor Crane: And is that a live possibility, or does that number — sixty-four out of a hundred — make the moderator explanation feel like a rescue?

Ben Okonkwo: Hm. I think it's both, and the ambiguity matters. But here's what the structural story adds that the moderator story can't: journals don't publish null results. Hiring committees don't reward replications. So even if some failures reflect missing moderators, the system has no mechanism for distinguishing those from genuine false positives — and no incentive to try.

Eleanor Crane: So the real question isn't whether individual papers are wrong — it's how science built a pipeline that can't tell the difference.

Ben Okonkwo: That's what today is actually about, I think — how a system gets built to reward exactly the behavior that breaks it.

Eleanor Crane: And that pipeline — I mean, the way it actually manufactures a false positive isn't dramatic at all. It's almost boring. Imagine someone buys twenty lottery tickets, wins once, and writes a paper titled 'I Have a System.' That's p-hacking in a sentence. You keep adjusting the analysis, shifting which variables you include, peeking at the results, until the number you need finally appears — and then you write up that version as if it were the only version you ever ran.

Ben Okonkwo: And you never mention the nineteen losing tickets.

Eleanor Crane: You never mention the nineteen losing tickets.

Ben Okonkwo: Now, there's a second move that makes this worse — actually, I'd say it's what turns a lucky result into a convincing-looking prediction. You run the analysis, you see what came out, and then you write the hypothesis as if you'd had it before you looked. HARKing — hypothesizing after results are known. It's like watching a dart hit a random spot on the wall and then walking up and drawing the bullseye around it. To any reader, it looks like you called your shot.

Eleanor Crane: Which is — wait, that's not even detectable from the outside, is it? The paper just reads as a clean prediction.

Ben Okonkwo: That's exactly the problem. And then — okay, compound it further. Run enough separate tests on the same dataset and a false positive becomes nearly guaranteed just by chance. Not by intent, not by bad faith. Mathematically inevitable. That's the multiple comparisons problem. The more questions you ask of the data, the more the data will eventually hand you something that looks like an answer.

Eleanor Crane: So the stated significance level — the number that's supposed to tell you how confident to be — is already optimistic before anyone does anything deliberately wrong.

Ben Okonkwo: Right — and when you stack p-hacking on top of HARKing on top of multiple comparisons, the false-positive rate inflation isn't a rounding error. It's the whole ballgame. Published 'significant' findings are carrying far less signal than the label says. And that's before a single result enters the file drawer and disappears because it didn't land.

Eleanor Crane: And the file drawer is where this compounds. Because it's not just that the false positives get published — it's that the null results, the ones that would balance the picture, they just... vanish. Someone runs the study, gets nothing, and puts it in a folder. The journal wouldn't take it anyway.

Ben Okonkwo: Right — and that asymmetry is what canonizes the error. Ego-depletion is the case I keep returning to. Roy Baumeister's paradigm. Earlier meta-analyses looked at the published literature and saw confirmation. Apparent consensus.

Eleanor Crane: A graduate student in 2010 reads those meta-analyses and thinks — the field has settled this. Builds a dissertation on it.

Ben Okonkwo: And the consensus was real, in the sense that it was in the literature. But the literature was — I mean, it was a biased sample. The publication-bias-corrected meta-analyses, starting around 2014, found no statistically reliable evidence for ego-depletion at all. Once you account for the file drawer, the effect disappears.

Eleanor Crane: The literature itself was lying. Not any single paper — the whole accumulated structure.

Ben Okonkwo: Now take that pattern into cancer biology, and the stakes shift completely. Begley and Ellis, 2012 — they tried to confirm fifty-three pre-clinical cancer studies. Landmark findings, top journals. Eleven percent replicated. Eleven.

Eleanor Crane: Wait — eleven percent of fifty-three. That's... six studies.

Ben Okonkwo: Roughly. And then the Reproducibility Project: Cancer Biology went further — fifty-three top cancer papers, 2010 to 2012, one hundred and ninety-three experiments total. Fifty experiments replicated. And not one of those original papers had fully described its experimental protocols. Seventy percent of replications required contacting authors directly just to get the reagents.

Eleanor Crane: Clinicians were reading those papers. That's the part I can't quite — and what comes next, whether the structural incentives can actually be reformed, that's even harder to sit with.

Ben Okonkwo: And that's the part that doesn't resolve cleanly — because the structural incentive isn't just indifferent to replication, it actively punishes it. Replication studies face publication bias against null results. There's no meaningful career reward for running one. The corrective mechanism has no professional engine.

Eleanor Crane: Which means the system can diagnose its own errors and still — what, shrug them off? That's not a bug someone forgot to fix.

Ben Okonkwo: Now, here's the sharpest version of this I've seen — there's economics PhD job market research showing that marginally significant results were associated with better academic placements than robust, solid ones. Better jobs. The researchers who were closest to the p-hacking line got rewarded for it.

Eleanor Crane: Wait — better placements for marginally significant results than for clean ones?

Ben Okonkwo: Better placements. Hiring committees, implicitly, valued the appearance of a surprising finding over a reliable one. That's the system working as designed — it's just designed wrong.

Eleanor Crane: So the question I keep turning over is — can you reform the mechanism without touching what the mechanism rewards? And people have tried. Registered Reports, which — I mean, PLOS Biology adopted this, Royal Society Open Science — the journal peer-reviews the design before you collect any data, provisionally accepts it, so the result can't be buried regardless of what comes out.

Ben Okonkwo: And the Center for Open Science built the whole scaffolding around this — Open Science Framework, the TOP Guidelines, pre-registration registries where you time-stamp your hypotheses before you run anything. In principle, it closes the HARKing loop entirely.

Eleanor Crane: In principle. But Campbell's Law is — actually, that's the critique that won't leave me alone. When a measure becomes a target, it loses its validity as a measure. Pre-registration becomes a badge. And badges get performed.

Ben Okonkwo: And Stroebe and Strack are still waiting in the wings — their moderator defense becomes the institutional antibody. Failed replication? Unmeasured contextual variable. Nothing to see. It insulates the original finding without engaging the evidence. Two layers of protection and the incentive structure never shifts.

Eleanor Crane: So if hiring committees still count Nature papers and journals still chase novelty — does pre-registration change the behavior, or just the paperwork?

Ben Okonkwo: Just the paperwork, I think. And that's — I mean, I can't shake the 36-out-of-100 number, not because it's a scandal about dishonest researchers, but because if you model it: journals reward novelty, hiring committees reward novelty, replication carries no professional engine — then 36 isn't shocking. It's almost the predicted output. A researcher can build a career on a novel false positive. Publishing a true negative costs years. Those aren't equivalent choices.

Eleanor Crane: And yet someone ran those replications. The Open Science Collaboration got a hundred psychologists to do work that had no career upside. That's the thing I can't quite place inside your model — not as a happy ending, but as a fact that doesn't fit the incentive story cleanly.

Ben Okonkwo: No, you're right. And I genuinely don't know what to do with that — whether it means the values and the incentives can diverge long enough to matter, or whether it's just a one-time coordination that won't hold without something structural underneath it.

Eleanor Crane: That's where I keep getting stuck too. What would have to change in what academia actually values — not what it requires on a checklist, but what it rewards when no one's watching — for accuracy to outrank surprise. I don't have an answer to that.

Ben Okonkwo: Neither do I. And I think that's probably the honest place this lands. Thanks for thinking through it with me.