Onpode
Cover art for Why journals and funders reward novel findings but scientists benefit from replication

Why journals and funders reward novel findings but scientists benefit from replication

August 24, 2026 · 13 min

Iris Holm & Lila Soto

The replication crisis in science is a structural problem: the Reproducibility Project: Psychology found only 36% of 97 high-profile studies replicated, yet replication work remains career-damaging because tenure and grants reward journal Impact Factor and citation counts — metrics that measure novelty, not accuracy.

Science operates under a structural incentive misalignment: individual researchers advance their careers primarily by publishing novel, statistically significant findings in high-prestige journals, attracting citations, and securing grants, while the long-term credibility of science as a system depends on whether those findings are reliable and replicable.

0:0012:30
Get the next episode on Science

Follow it free — new episodes land in your feed.

Or make your own — any topic, in minutes

More Onpode episodes on Science

About this episode

A researcher does careful, rigorous work — replicates a famous study, finds it doesn't hold — and the paper goes nowhere. Meanwhile the original has five hundred citations and a tenure case built on top of it. That's not a failure of individual integrity. It's the reward structure doing exactly what it was designed to do. This episode works through why science's incentive architecture selects against the one activity — replication — that would actually stress-test the literature. It starts with the numbers: the Reproducibility Project found only 36% of 97 high-profile psychology studies held up, and surveys of more than 1,500 scientists found over 70% had already personally failed to reproduce someone else's results. That's not a fringe problem. From there it traces the mechanism: Journal Impact Factor measures citation frequency, not accuracy. Universities built hiring and tenure decisions on that metric. The metric measures buzz, and buzz became the definition of quality. A replication that overturns a finding is, by construction, less citable than the original — so the correction is always worth less than the error it corrects. The episode also takes seriously the tools that exist to fix this — preregistration, Registered Reports, the Center for Open Science — and asks honestly why adoption stays uneven. The answer keeps coming back to the same room: the hiring meeting, the tenure vote, the grant review. Until those rooms change what they count, the 36% doesn't move. No cleaner version of that exists.

Frequently asked

What percentage of psychology studies replicated in the Reproducibility Project?

The Reproducibility Project: Psychology, led by Brian Nosek and the Center for Open Science, found that only 36% of 97 high-profile psychology studies successfully replicated. The studies tested were not fringe work — they were prominent, widely cited findings in the field.

What is p-hacking and why do scientists do it?

P-hacking is slicing or reanalyzing data in multiple ways until a result clears the p < 0.05 significance threshold. Scientists do it because publication bias means null results rarely get published, and without publications, postdocs lose funding and researchers lose jobs. It is individually rational behavior inside a broken reward structure.

Why don't replication studies get published or cited?

A replication study that overturns a high-profile finding is structurally worth less than the original, because the original already has years of downstream papers citing it. The corrective paper gets cited mainly by others writing about corrections — making it nearly invisible under Impact Factor and citation-count metrics used in hiring.

Do preregistration and Registered Reports fix the replication crisis?

Preregistration and Registered Reports — where journals accept papers based on methodology before results exist — can close the p-hacking opportunity and attack publication bias at the journal level. But adoption is voluntary and uneven. Tenure committees still count Nature papers, not preregistration metadata, so the core incentive remains unchanged.

Why do replication rates differ so much between psychology and economics?

Psychology replication rates sit around 36%, while economics and social science reach 60–70%, despite both fields sharing the same publish-or-perish incentive structure. A leading explanation is data architecture: economists typically use large administrative datasets where p-hacking is harder to execute, while many psychology studies use small lab samples of around 60 participants.

Grounded in 12 sources
Beyond Citations: Comparing Scholarly, Policy, and Patent Impact Across the FT50 Journals · arxiv.org
Ranking journals: Could Google Scholar Metrics be an alternative to Journal Citation Reports and Scimago Journal Rank? · arxiv.org
Does the use of open, non-anonymous peer review in scholarly publishing introduce bias? Evidence from the F1000 post-publication open peer review publishing model · arxiv.org
Operationalizing Narrative Entropy (Sn): A Two-Scene Registered Pilot Report and Pre-Validation Protocol · arxiv.org
The influence of publication ranking specifications on publication strategy and academic careers in business administration · pmc.ncbi.nlm.nih.gov
Preregistration: A Key to Credible Real‐World Evidence Generation · pmc.ncbi.nlm.nih.gov
The misalignment of incentives in academic publishing and ... · pnas.org
(PDF) Philosophy of Science and The Replicability Crisis · researchgate.net
Preserving scientific integrity in academic publishing: Navigating artificial intelligence, journal policies, and the impact factor as a quality indicator · sciencedirect.com
Incentives and the replication crisis in social sciences: A critical review of open science practices · sciencedirect.com
The Replication Crisis in Science: Causes and Solutions | Adjmal Sarwary · adjmal.com
Psychology Replication Crisis: What Happened — CASRAI · casrai.org
Read transcript

Lila Soto: Okay, I have to tell you the thing that's been nagging me since I went down this rabbit hole — there's a researcher, doing careful, rigorous work, replicates a famous study, finds it doesn't hold. Solid methodology. And then the paper goes nowhere. Meanwhile the original finding has, like, five hundred citations.

Iris Holm: And a grant.

Lila Soto: Right — and a grant, and a tenure case. And what I keep wondering is whether that's a flaw in the system or just... the system working exactly as designed.

Iris Holm: It's the system working as designed. That's the uncomfortable part.

Lila Soto: Which is what today is actually about — this incentive misalignment between what rewards individual scientists and what would actually benefit science as a whole. And the place where that becomes concrete is the Reproducibility Project: Psychology. Brian Nosek, the Center for Open Science, 97 studies — and only 36% replicated.

Iris Holm: Thirty-six. Not thirty-six percent of a fringe. Ninety-seven high-profile studies.

Lila Soto: Hm, and what gets me is — this isn't even a secret. Surveys of more than 1,500 scientists found over 70% had already personally failed to reproduce someone else's results. It's the normal experience. And we still treat replication work like it's career poison.

Iris Holm: Because it is. That's not rhetoric — the publish-or-perish norm structurally punishes the one activity that would actually stress-test the literature.

Lila Soto: Yeah, so the question we're really trying to work out — why does a system built by people trying to do good science end up selecting against reliable science?

Iris Holm: Good faith, bad architecture.

Lila Soto: Good faith, bad architecture — yeah, that's exactly the frame.

Iris Holm: Now: what does it take to change the architecture when every institution that would have to change it is also winning under the current rules?

Lila Soto: But that's the part I keep snagging on — like, how did the architecture get wired that way in the first place? Because someone had to decide that Journal Impact Factor meant something.

Iris Holm: Think of it like a movie review site where the score only counts how many people clicked the review. Not whether the movie was good. Just clicks. Now imagine studios started greenlighting movies purely to get clicked.

Lila Soto: Oh. That's — yeah, that's exactly it.

Iris Holm: Impact Factor counts citation frequency. Not accuracy. And then universities took that number and built tenure decisions on top of it. Hiring decisions. Grant allocations. The metric measures buzz, and we made buzz the definition of quality.

Lila Soto: And it doesn't stop at the journal level — I mean, the same logic runs through h-index, citation count. All the bibliometric indicators are doing the same work at the individual researcher level. So it's kind of... stacked.

Iris Holm: Stacked is the right word. And here's the mechanism nobody says plainly: a replication study that overturns a high-profile finding is, by construction, less citable than the original. Because the original has five years of downstream papers all citing it.

Lila Soto: Wait — so the corrective is structurally worth less than the error it corrects?

Iris Holm: Always. The original gets cited by everyone who built on it. The correction gets cited by... people writing about corrections. PNAS published research making exactly this point — framing the whole thing as systemic, not moral. It's not that scientists are lying. The market design produces this outcome from honest behavior.

Lila Soto: Which is almost worse? Like, if it were just bad actors you could fix it with enforcement. But if it's rational behavior inside a broken structure — yeah, I don't know what the lever is.

Iris Holm: The lever is what actually determines whether someone gets hired. And right now that's Impact Factor and citation count. Preregistration doesn't touch that. Neither does open peer review.

Lila Soto: Hm. So top journals preferentially publish the highly novel study — the one that's never been done before — even if that novelty makes it harder to verify. And that study lands the citations, which feeds the h-index, which feeds the tenure case. The whole chain runs on novelty.

Iris Holm: And the researcher who does careful replication work — solid methodology, null result — gets rejected everywhere, struggles to fund the next project. The question is whether any institution is actually willing to change what it rewards at the moment that counts. The hiring meeting. The tenure vote.

Lila Soto: And that's where it stops being abstract — because she's not sitting there plotting. She's just doing the math on her own survival. The postdoc, eight months left on her contract, null result in front of her. What does the system train her to do next?

Iris Holm: Run the subgroups.

Lila Soto: Run the subgroups. Try a different control variable. Slice the data a different way until something clears p less than 0.05. And that's p-hacking — not fabrication, just... iteration until the number cooperates.

Iris Holm: And individually? Totally rational. Publication bias means null results don't get published. No publications, no renewal. The math is clear.

Lila Soto: I mean — it's not even, like, a moral failure. The structure produced this outcome. She didn't invent the incentive.

Iris Holm: Right. And then the companion move — once the data shows something unexpected, you write the introduction as though you predicted it. That's HARKing. Hypothesizing After the Results are Known. Exploratory finding, dressed as confirmatory.

Lila Soto: Wait — so the paper reads like a hypothesis test, but it's actually pattern-matching on the same dataset that generated the pattern.

Iris Holm: Every time. And publication bias closes the loop — because the journals are only seeing the sliced version, the one that cleared 0.05. The null result stays in a drawer. So the literature fills up with findings that look solid and aren't.

Lila Soto: Which is actually — okay, I want to say this carefully — that's a market-design failure, not a misconduct problem. The reward structure made this individually rational and collectively self-destructive. Like, she did what the system told her to do, and the system got worse.

Iris Holm: If you enforce against individuals without changing the reward structure, you just get more careful p-hacking. The behavior is load-bearing. It's holding the career up.

Lila Soto: Hm. And what nobody talks about enough — she probably knows. Like, she's not confused about what she's doing. She's made a calculation.

Iris Holm: Which is the bleaker version of the problem.

Lila Soto: Yeah. And the numbers that actually show how deep this goes — the Reproducibility Project, what the replication rates look like across fields, what preregistration and Registered Reports can actually fix and what they can't — that's kind of where this gets even harder to sit with.

Iris Holm: But here's where the numbers get weird — because if p-hacking is rational everywhere, the replication rates shouldn't look this different across fields. Psychology: 36%. Economics and social science: 60 to 70. Same broken incentive structure. Very different outcomes.

Lila Soto: Wait — that's a huge gap. Why does economics look so much cleaner?

Iris Holm: Nobody's settled it. But my read — economists use larger datasets almost by default, administrative records, census data. Harder to slice into significance when n is a hundred thousand. Psychology runs a lot of sixty-person lab studies where the wiggle room is enormous.

Lila Soto: So it's not that economists are more virtuous — it's that the data structure makes p-hacking harder to execute cleanly.

Iris Holm: Which changes the whole diagnosis. If this were purely a culture problem, you'd fix it with culture. But if method architecture drives the number — then the fix has to be structural too.

Lila Soto: And then there's preclinical cancer biology sitting at around 40%, which is — I mean, that's not a methodology seminar problem. That's researchers building drug trials on foundations that don't hold half the time.

Iris Holm: 40% in preclinical work, yes. And that's where the abstraction ends.

Lila Soto: Like — a compound gets flagged as promising because the lab result looked solid. It enters trials. The lab result was one of the 60% that didn't hold. And you're asking patients to be in that trial.

Iris Holm: Frankly, that's the case for why the reform tools matter. Preregistration — you register your hypothesis, your design, your analysis plan on the Open Science Framework before you touch the data. The p-hacking opportunity mostly closes.

Lila Soto: And Registered Reports go further, right? The journal accepts or rejects based on methodology before results exist. In-Principle Acceptance. So the outcome literally cannot determine whether it gets published.

Iris Holm: That's the mechanism. You decouple the publication decision from the result. Which is — look, that's the only intervention that actually attacks publication bias at the journal level, not just the researcher level.

Lila Soto: Hm. But they're voluntary. The Open Science Framework is there, the infrastructure exists — and adoption is still uneven.

Iris Holm: Voluntary and uneven is the whole problem. A tenure committee reads a CV. They count publications in high-Impact-Factor journals. Preregistration is metadata. It doesn't appear in that count.

Lila Soto: So Registered Reports end up as a parallel track — the methodologically careful researchers use them, and the prestige machine keeps running on novelty beside them. And I kind of wonder if that actually makes it worse — like it gives the system a safety valve without fixing the pressure.

Iris Holm: The question is whether a tenure committee will ever count a Registered Report the same as a Nature paper. And the answer depends entirely on whether funders and hiring bodies enforce it — not whether the journals offer it.

Lila Soto: So the tools exist. The enforcement moment — the hiring meeting, the grant review — that's still running the old criteria. And until those rooms change what they reward, the 36% doesn't move.

Iris Holm: And that's the thing nobody in the reform conversation wants to say out loud. The Open Science Framework exists. Preregistration exists. Registered Reports exist. The architecture for a better system is sitting there. And a tenure committee still opens a CV and counts Nature papers. That's it. That's the whole enforcement problem in one room.

Lila Soto: Yeah. And I guess what I keep wondering — not rhetorically, like actually — what would it take for a university to genuinely hire someone whose whole career is careful replication work? Not one replication study among other things. A whole career of it. What does that candidate even look like to a hiring committee running on Journal Impact Factor?

Iris Holm: Invisible. Or — I mean, worse than invisible. They look like someone who couldn't generate their own ideas. That's the read. Even if the work is the most scientifically valuable thing produced in that department that decade.

Lila Soto: Which is kind of where I land, actually. Not with an answer — just with the weight of it. The reforms are real. Preregistration is real, Registered Reports are real, Brian Nosek built actual infrastructure at the Center for Open Science. And it's still not enough if the room where it counts never changes what it's counting.

Iris Holm: That's the uncomfortable place this leaves you. Good tools. Broken adoption. The 36% doesn't move until the hiring meeting does. I don't have a cleaner version of that.