Onpode
Cover art for Why the peer-review system rewards novelty over accuracy and replication

Why the peer-review system rewards novelty over accuracy and replication

August 24, 2026 · 12 min

Sarah Lin & Dr. Nathan Hayes

82% of researchers never submit null results — not because journals reject them, but because they self-censor first. Yet 51% of null-result papers that do get submitted are accepted. This gap reveals that peer review's publication bias is driven by incentive architecture, not gatekeeping: novelty attracts citations, so null findings simply offer less.

The peer review system in academic publishing systematically favors novel, statistically significant findings over null results and replications—a phenomenon documented as "publication bias" or the "file drawer problem."

0:0012:16
Get the next episode on Science

Follow it free — new episodes land in your feed.

Or make your own — any topic, in minutes

More Onpode episodes on Science

About this episode

The peer-review system wasn't designed to fail. That's what makes it hard to fix. This episode works through the specific mechanics of why the scientific literature ends up as a systematically inflated positive sample — and why the people best positioned to change it are often the ones it's already working for. The episode opens with a number that reframes the whole problem: 82% of researchers with null results never submit them, while 51% of the null-result papers that do get submitted are accepted. The gap between those two figures is where the file drawer problem actually lives — not in hostile gatekeepers, but in a rational career calculation made long before any editor sees anything. From there, it traces how novelty bias operates not as individual preference but as the output of an asymmetric reputational structure: no penalty for a false positive that quietly fails to replicate, real visible cost for passing on the paper that becomes a landmark. It connects that structure to p-hacking and HARKing — neither of which requires bad intentions, both of which are incentive-compatible responses to the same pressure. The episode takes Registered Reports seriously as a reform and then asks whether they reach far enough — including a finding from F1000Research's open peer-review experiment, where bias didn't disappear, it just relocated. And it ends on a harder problem: even if every structural reform worked, the statistical tools underneath may still be misspecified for the inferences we're asking them to make. No villain, no easy resolution. Just a clear account of a machine working exactly as its incentives specify.

Frequently asked

What is the file drawer problem in peer review?

The file drawer problem is the systematic non-publication of null results, failed replications, and inconclusive findings. Researchers never submit them — not because journals reject them, but because novelty bias means accepted null results still circulate less and advance careers less. The published literature becomes an unrepresentative, inflated positive sample.

Why do journals favor novelty over replication?

Journals favor novelty because editors face asymmetric reputational incentives: publishing a false positive that later fails to replicate carries no formal penalty, but rejecting a manuscript that becomes a landmark is a lasting mark on an editor's record. That asymmetry — not individual bias — systematically tilts acceptance toward surprising, counterintuitive findings.

What is p-hacking and why does it happen?

P-hacking is the practice of adjusting analytic choices — which variables to include, when to stop collecting data, which statistical test to run — until a p-value drops below 0.05. Each individual choice is defensible, which is why p-hacking is better understood as incentive-compatible behavior than misconduct: a p below 0.05 is the ticket to publication.

Do Registered Reports actually fix publication bias?

Registered Reports — where peer review and journal commitment happen before data collection — structurally break the link between result and acceptance, targeting the core mechanism of publication bias. However, if citation counts for preregistered null results do not rise, the prestige gradient survives the reform, and upstream funding decisions that still reward novelty remain untouched.

What did the replication crisis reveal about the scientific literature?

The replication crisis revealed that a substantial proportion of published, statistically significant findings — particularly in psychology — fail when independent researchers attempt to reproduce them. This confirmed that decades of publication bias had produced a literature where both the prevalence and magnitude of true effects were systematically overstated, misleading clinical guidelines and funding priorities built on that record.

Grounded in 11 sources
Does the use of open, non-anonymous peer review in scholarly publishing introduce bias? Evidence from the F1000 post-publication open peer review publishing model · arxiv.org
Publication Bias (The "File-Drawer Problem") in Scientific Inference · arxiv.org
Reforming research funding: Combining editorial preregistration with grant peer review · arxiv.org
How to Tell When a Result Will Replicate: Significance and Replication in Distributional Null Hypothesis Tests · arxiv.org
Estimating publication bias in meta-analyses of peer-reviewed studies: A meta-meta-analysis across disciplines and journal tiers · doi.org
Publication bias in experimental philosophy: a survey on the file drawer problem, research attrition, and open science practices · doi.org
Publication bias and the failure of replication in experimental psychology · link.springer.com
Replication Research, Publication Bias, and Applied Behavior Analysis - PMC · pmc.ncbi.nlm.nih.gov
Publication bias and the chase for statistical significance - PMC · pmc.ncbi.nlm.nih.gov
Incentives for Research Effort: An Evolutionary Model of Publication Markets with Double-Blind and Open Review · pmc.ncbi.nlm.nih.gov
Scientific Utopia: II. Restructuring Incentives and Practices to Promote Truth Over Publishability · pmc.ncbi.nlm.nih.gov
Read transcript

Sarah Lin: Nathan, I want to start with a thought experiment — imagine you've just run a study and the result is null. No effect. Clean, well-powered, null. What do you do with it?

Dr. Nathan Hayes: Honestly? The correct answer is: submit it. The likely answer, given the incentive structure, is: put it in a drawer.

Sarah Lin: And that — that drawer — is the whole thing we're here to talk about. Because eighty-two percent of researchers are apparently choosing the drawer. Not because the journals reject them, but because... they never find out. They don't submit.

Dr. Nathan Hayes: Wait — you're saying the self-censorship happens before any institutional rejection.

Sarah Lin: Before. And meanwhile, more than half of null-result papers that do get submitted are accepted. Fifty-one percent. So the anticipated rejection isn't matching the actual one.

Dr. Nathan Hayes: That gap — eighty-two percent never submitting, fifty-one percent accepted when they do — that's where the file drawer problem actually lives. It's not a story about hostile gatekeepers. It's a story about a belief that has become so load-bearing that the evidence against it doesn't dislodge it.

Sarah Lin: Right — but is the belief wrong, or is it reading something true that the acceptance rate doesn't capture?

Dr. Nathan Hayes: Now that's the important distinction. The belief is correctly reading the incentive, just not the acceptance probability. Novelty bias is real — editors and reviewers prefer surprising, counterintuitive findings because those attract citations. A null result, even accepted, doesn't circulate the same way. So the career calculus still tilts against submission.

Sarah Lin: So the file drawer fills up not because the door is locked but because the room on the other side... offers less.

Dr. Nathan Hayes: And what accumulates in that drawer is precisely what the replication crisis revealed was missing — the published literature ends up as a systematically inflated positive sample, which is publication bias in its formal definition. And then when independent researchers try to reproduce those published findings, a substantial share fail. The record was never representative.

Sarah Lin: And no one in that chain was — I mean, there's no villain here. Editors, reviewers, researchers — they're all responding rationally to the same asymmetry.

Dr. Nathan Hayes: The asymmetric reputational incentives guarantee it — publish a false positive that later fails and you face no formal penalty; reject a manuscript that later proves influential and that's a visible mark on your record. The structure produces the distortion. The individuals are almost incidental.

Sarah Lin: But that's almost the clean version, right? Like — the individuals are incidental, the structure does it, fine. But what does the structure actually look like from inside an editorial decision?

Dr. Nathan Hayes: Imagine a publication that prints restaurant reviews. They run a glowing review of a place that turns out to be mediocre. Nobody remembers. The clip doesn't follow them. But if they pass on a restaurant that goes on to win three Michelin stars, that rejection gets cited forever. So of course they print the glowing reviews. That is the whole structure.

Sarah Lin: So we're not looking for bad actors. We're watching good actors make the move that the board rewards.

Dr. Nathan Hayes: Exactly that. And the word I'd use is incentive-compatible — meaning, doing what the system rewards isn't just tempting, it's the rational move. An editor who favors novelty isn't corrupt. They're reading the asymmetric reputational incentives correctly. No penalty for publishing a false positive that fails to replicate. Real, visible reputational cost for rejecting the manuscript that becomes a landmark.

Sarah Lin: Wait — and citation metrics are how that cost actually lands? Like, it's tracked?

Dr. Nathan Hayes: Citation metrics and historical narrative, yes. Those are the concrete mechanisms. Someone writes the history of the field and they note — this journal passed on that paper. That becomes the editor's record. Whereas the false positive that quietly never replicates? It just fades. It doesn't get a named moment.

Sarah Lin: So novelty bias isn't even a preference, really. It's — um — it's the output of that asymmetry being applied consistently over time.

Dr. Nathan Hayes: Right — and now the harder part to sit with. A well-meaning editor, operating completely rationally inside this system, would produce the same distortion as a biased one. The incentive is compatible with good intentions. That's what makes it — actually, this is the part I don't think most reform proposals reckon with — you can't fix it by finding better people.

Sarah Lin: Oh. Yeah.

Dr. Nathan Hayes: Because the novelty bias isn't downstream of individual preference — it's upstream of it. It shapes what gets framed as interesting before any single person makes a call.

Sarah Lin: Which means — and this is the part I'm sort of sitting with — the distortion happens before the file drawer even opens. It happens when the researcher decides what question to ask in the first place, because they already know what the room rewards.

Dr. Nathan Hayes: And that's where publication bias bleeds into something harder to measure — not just what gets published, but what gets designed. What gets attempted. That's a different problem than a drawer full of null results. That's a literature that's missing questions it never thought to ask.

Sarah Lin: And if the literature is missing questions it never asked... I mean, that's what the file drawer problem actually costs, right? It's not just the results that disappear. It's the whole shape of what we think we know.

Dr. Nathan Hayes: The formal definition matters here — the file drawer problem is specifically that null results, failed replications, inconclusive findings, they never get submitted, or they get rejected, and they sit. Unpublished. What's left standing is an unrepresentative positive sample dressed up as the scientific record.

Sarah Lin: Can I make it specific for a second?

Dr. Nathan Hayes: Please.

Sarah Lin: There's a graduate student — Thursday evening, let's say — and they've just finished running an experiment they built around a published effect size that looked significant. They ran it. They got nothing. And now they're sitting with that null result, and... um... they're doing the math. Not the statistical math. The career math. And they quietly don't write it up.

Dr. Nathan Hayes: And that single quiet decision — that's one more result added to the file drawer. Not with any bad intention. Just... rational.

Sarah Lin: And the published effect size they built their whole design around? Still sitting in the literature. Unchallenged. Because the challenge is in the drawer.

Dr. Nathan Hayes: Which is exactly what the replication crisis surfaced — when independent researchers actually tried to reproduce published, statistically significant findings in psychology and other fields, a substantial proportion failed. That's the empirical proof that the accumulation happened. The drawer was real and it was full.

Sarah Lin: And not just full — systematically full. Like, it's not random noise that got suppressed. It's a specific direction.

Dr. Nathan Hayes: Correct. Publication bias means the aggregate literature overstates both the prevalence and the magnitude of true effects. The effect sizes in the record are inflated. The rates of significance are inflated. So when a clinician builds a treatment protocol on that record, or a funding body prioritizes one research direction over another — they're working from a tilted sample and they don't know it.

Sarah Lin: Oh, that's — yeah. The clinical guidelines. We built those on this.

Dr. Nathan Hayes: Interventions, follow-up studies, policy decisions — all downstream of a literature that was never representative. And what makes it harder is — actually, the part we haven't touched yet is why individual researchers under that pressure start adapting their analysis choices in ways that are also rational, and why that makes even the accepted papers unreliable in ways that structural reforms like Registered Reports may not fully reach.

Sarah Lin: So the drawer isn't even the whole wound. The stuff that does get published... some of it is also kind of shaped by the pressure to get there.

Dr. Nathan Hayes: Shaped, yes — and that's where p-hacking comes in. Which sounds like misconduct but is actually just... exploiting degrees of freedom. You have choices — which variables to include, which subgroup to analyze, when to stop collecting data, which statistical test to run. And you keep making those choices until p drops below 0.05.

Sarah Lin: And p below 0.05 is the ticket.

Dr. Nathan Hayes: The ticket, yes. And none of those individual choices are obviously wrong — they're all defensible. That's what makes it — I mean, you can't point at a single decision and say fraud. The researcher isn't lying exactly. They're navigating.

Sarah Lin: And then there's HARKing. Hypothesizing After Results are Known. Which is — um — you find the significant result and then you write the hypothesis as though you'd predicted it from the start.

Dr. Nathan Hayes: Presenting exploratory work as confirmatory. Which inflates confidence dramatically — a hypothesis that was actually generated by the data appears to be independently testing it.

Sarah Lin: So both of those — p-hacking, HARKing — they're not the bad actor moves. They're the rational moves.

Dr. Nathan Hayes: Incentive-compatible. Which is exactly why Registered Reports are the most structurally serious reform on the table — peer review happens before data collection, the journal commits to publish regardless of outcome. You break the link between result and acceptance.

Sarah Lin: That actually... targets the mechanism. Not just the symptom.

Dr. Nathan Hayes: It does. And I want to give it that. But — okay, so F1000Research ran an experiment with post-publication open peer review, non-anonymous, reviewers named publicly. The hypothesis was that transparency kills bias. What they actually found was new biases: deference to whatever the first visible review said, and country-affiliation effects — where a reviewer's institutional location influenced reception. Bias didn't disappear. It relocated.

Sarah Lin: Oh — it just found a different shape.

Dr. Nathan Hayes: Which is why Lutz Bornmann and Gerald Schweiger argued the reform has to reach upstream — integrate preregistration with grant peer review, not just publication. Because if the funding decision still rewards novelty, the incentive survives the editorial reform completely intact.

Sarah Lin: So the publication fix is downstream of where the pressure actually starts.

Dr. Nathan Hayes: And now — actually, this is the part that doesn't fit anywhere clean — Fintan Costello and Paul Watts showed that standard significance tests only analyze variation within a single experiment. Between-experiment variation — the spread across independent replications — that's just not in the test. So even a perfectly preregistered, fully transparent study is producing a confidence interval that systematically understates uncertainty. The statistical tool itself is misspecified for the inference we're asking it to make.

Sarah Lin: Wait — so even if every reform worked, the math underneath is still giving us the wrong answer?

Dr. Nathan Hayes: The asymmetric incentive architecture and the statistical misspecification are separate problems. You could fix publication bias entirely and still have a replication crisis — because a single significant result was never as informative as the p-value made it look. The reforms are real. They're just not — they're not touching that layer.

Sarah Lin: The thing I keep not being able to get past — and maybe there's no clean way through it — is that the people who could actually change this are the people it's already working for. The editors whose reputations are built on the bold, surprising paper. The researchers whose careers moved forward because novelty bias rewarded them. I mean, asymmetric incentives persist because the system is, for certain people, not broken at all.

Dr. Nathan Hayes: That's genuinely where I can't resolve it either. Registered Reports are real. The logic is sound. But the citation counts for preregistered null results — I haven't seen evidence they circulate differently than traditional negatives. And if they don't, the prestige gradient survives the reform. You've added process. You haven't moved the reward.

Sarah Lin: Which is — um — I think what sits hardest with me. It's not a broken machine. It's a machine working exactly as its incentives specify.

Dr. Nathan Hayes: And we talk about reforming peer review as though the output is a malfunction. The file drawer problem, the replication crisis, novelty bias — they're not malfunctions. They're the designed output, given what the system is actually optimized for. The question is only whether we want different outputs badly enough to change what it's optimizing for. And I don't know who 'we' even is in that sentence.

Sarah Lin: Yeah. I don't either. That graduate student on Thursday evening — she's not going to reform it. And the editor who passed on the paper that became a landmark, they learned something, but not something that changes the structure. It's — I don't have a resolution. I just have the question sitting there.

Dr. Nathan Hayes: That's an honest place to be with it. Thank you for thinking through it alongside me.