Onpode
Cover art for How publication incentives systematically bias toward surprising results over robust ones

How publication incentives systematically bias toward surprising results over robust ones

August 2, 2026 · 15 min

Iris Holm & Cyrus Reed

Publication bias has been documented since Sterling audited psychology journals in 1959 and found nearly all published results were statistically significant — yet the pattern held through 1995 and beyond. The core problem is pre-submission suppression: researchers quietly shelve null results before editors ever see them, a career calculation compounding across thousands of scientists.

The peer-review and publication system exhibits a durable structural bias favoring novel, statistically significant findings over effect size, replicability, and null results. This pattern has been documented across disciplines for decades. Robert Rosenthal described the "file drawer problem" in 1979, noting that unpublished null results distort the apparent literature.

0:0014:37
Get the next episode on Science

Follow it free — new episodes land in your feed.

Or make your own — any topic, in minutes

More Onpode episodes on Science

About this episode

In 1959, a researcher named Sterling audited psychology journals and found nearly every published study reported a statistically significant result. He repeated the audit in 1995. The numbers were the same. This episode works through why — and why sixty-five years of knowing about publication bias has done almost nothing to correct it. The argument isn't that researchers are dishonest or that journal editors are corrupt. It's that the system selects for false positives the way evolution selects for traits: not through intention, but through structure. Small, underpowered studies are actually more likely to clear the p<0.05 threshold by luck, which makes them more likely to be submitted, more likely to be published, and more likely to become foundational — until a replication attempt that will itself struggle to get published tries to undo the damage. The episode traces how three independent nodes — publish-or-perish pressure, journal impact factor, and funder track-record requirements — each apply the same novelty premium without coordinating. It looks at why Registered Reports, which commit journals to publishing based on study design rather than outcome, are a genuine improvement that still hasn't cracked the underlying incentive problem. And it ends honestly: the fix has to reach what gets rewarded at tenure time, and nothing described so far actually changes the Tuesday-morning calculation for a researcher sitting with a null result and a contract to worry about.

Frequently asked

When was publication bias first documented in scientific research?

Publication bias in academic research was documented at least as early as 1959, when Sterling audited psychology journals and found nearly all published papers reported statistically significant results. Sterling confirmed the same pattern in 1995 — thirty-six years later with no measurable change in the skew of the published literature.

What is the file drawer problem in research?

The file drawer problem, named by Robert Rosenthal in 1979, describes how null or negative results never get submitted for publication. Researchers calculating career costs simply open a new experiment instead. The pre-submission gap accounts for roughly sixty percentage points of the positive-result skew — far more than editorial rejection of submitted null results.

What is p-hacking and is it the same as fraud?

P-hacking is the practice of making small, individually defensible analytical choices — excluding outliers, trying different model specifications — until a result clears the p less than 0.05 threshold. It is not fraud; no data is fabricated. The problem is that none of these choices are disclosed, so a paper that was exploratory reads as confirmatory.

What did the 2015 Open Science Collaboration replication study find?

The Open Science Collaboration's 2015 replication project attempted to reproduce 100 published psychology studies and found that only 36 to 39 percent held up. The result illustrated that a literature built on small, underpowered studies with borderline p-values — the kind the publication incentive system selects for — does not reliably reflect real effects.

Do Registered Reports fix the replication crisis?

Registered Reports — where peer review and journal commitment happen before data collection — do increase null-result publication and reduce outcome-switching. But they leave the career incentive untouched: journal impact factor rewards novel positive results through citation counts, so a carefully conducted null finding still contributes little to a tenure file.

Grounded in 11 sources
Pre-registration for Predictive Modeling · arxiv.org
Stakeholder views on publication bias in health services research · doi.org
Registered reports: prospective peer review emphasizes science over spin. · doi.org
Current Incentives for Scientists Lead to Underpowered Studies with Erroneous Conclusions | PLOS Biology · journals.plos.org
Effect size, not sample size, predicts the replicability of ... · journals.plos.org
Detecting and avoiding likely false-positive findings – a practical guide · onlinelibrary.wiley.com
Reproducibility of Scientific Results · plato.stanford.edu
The misalignment of incentives in academic publishing and implications for journal reform · pmc.ncbi.nlm.nih.gov
The replication crisis has led to positive structural, procedural, and community changes · pmc.ncbi.nlm.nih.gov
A framework for assessing the trustworthiness of scientific research findings1 | PNAS · pnas.org
Replication Crisis - an overview · sciencedirect.com
Read transcript

Cyrus Reed: Iris, tell me if this is a weird way to start — I've been thinking about Tuesday mornings.

Iris Holm: That is a weird way to start.

Cyrus Reed: No but stay with me — a researcher, Tuesday morning, staring at data that shows nothing. No effect. And they just... decide not to write it up. Not because anyone told them not to, but because they've internalized that null results aren't worth submitting. That decision, multiplied across thousands of scientists, is actually — wait, that might be the whole story. Like, the bias isn't even at the journal level.

Iris Holm: Pre-submission suppression. The study that gets published is the one with twenty participants, a p-value of 0.049, no replication — but it's the one that made it out of someone's desk drawer.

Cyrus Reed: And becomes foundational! Someone builds a career citing that study, and then the Open Science Collaboration in 2015 tries to replicate a hundred studies like it — and only thirty-six to thirty-nine percent hold up. Three-sixths of a century after Sterling first audited psychology journals and found the same skew in 1959.

Iris Holm: Hold on — you're saying Sterling saw this in 1959.

Cyrus Reed: Nearly all published papers reported statistically significant results. 1959. So the replication crisis — the name feels new, the thing is ancient.

Iris Holm: That's what we're here to work out today. Publication bias, the replication crisis — where it comes from, why it persists. And the question underneath all of it: if the system has known for sixty-five years, what is actually keeping it broken?

Cyrus Reed: And my hunch — which I want to test — is that the more interesting culprit isn't a bad journal editor. It's the structure of what gets rewarded. But I don't know if that makes it fixable or just... depressing.

Iris Holm: Probably both.

Cyrus Reed: Probably both — okay so the underpowered study getting published because it was underpowered, not despite it, that's the entry point. Because once I understood that mechanism it reframed everything else.

Iris Holm: Then let's start there. Small sample, lucky p-value, landmark paper — and the larger replication that should have killed it just disappears. That asymmetry is the whole engine.

Cyrus Reed: But here's where I want to push on that — because 'lucky p-value' makes it sound like the researcher got unlucky or something passive happened. The weirder truth is that the researcher *made* the luck. Imagine a grad student, Friday afternoon, data's in, p-value comes out at 0.07. Not significant. And she thinks — wait, are those two data points actually outliers? Plausibly, right? She excludes them. Now it's 0.04. That's the version that goes in the paper. And she didn't lie. She didn't fabricate anything.

Iris Holm: Right — and that's researcher degrees of freedom in one Tuesday-afternoon moment.

Cyrus Reed: Friday afternoon, but — yeah, exactly. And the thing is, she probably didn't even decide to exclude them *because* of the p-value. She found a reason that felt real. That's the mechanism. It's not fraud, it's not even conscious, it's just... the choices were always there and the incentive pointed one direction.

Iris Holm: Darts, blindfolded. Infinite throws. You only show the board the ones that hit.

Cyrus Reed: That's — yeah. That's p-hacking. Not one throw, not cheating, just... selective showing.

Iris Holm: But isn't that just — doing statistics? Like, you're *supposed* to check your data for outliers.

Cyrus Reed: That is *exactly* the question — and it's actually the hard part. Because yes, you should check for outliers, and yes, you should try different model specifications. The problem isn't any single choice. It's that none of those choices get disclosed. So the reader sees a clean confirmatory result, but what actually happened was — wait, this is the HARKing piece — she probably wrote the hypothesis *around* the result she found. Hypothesizing After Results are Known. The paper reads like prediction; it was actually exploration.

Iris Holm: And that's not detectable from the paper.

Cyrus Reed: Not even slightly. And here's what that does at scale — small, underpowered studies, twenty participants, borderline everything, they're actually *more* likely to clear 0.05 by luck. So they're more likely to get submitted. More likely to get published. Ioannidis put the framework for this in PLOS Medicine in 2005 — 'Why Most Published Research Findings Are False' — and the math is brutal. The system doesn't just tolerate this, it selects for it.

Iris Holm: Hold on — 'selects for it' is doing a lot of work there. Spell that out.

Cyrus Reed: Okay — so statistical significance at p less than 0.05 tells you the result is unlikely under the null. That's it. It says nothing about whether the effect is *large* or *real* or *meaningful*. But journals treat it like a publication criterion anyway. Effect size — how big the thing actually is — basically doesn't enter the gate decision. So a tiny, noisy effect that squeaked past 0.05 gets the same door as a robust one. The small underpowered study is more likely to produce that squeak by luck. Robert Rosenthal was already pointing at this architecture in 1979 when he named the file drawer problem.

Iris Holm: So p less than 0.05 was never supposed to be a one-number filter, and yet.

Cyrus Reed: And yet! That's the whole thing — it functions as one. And that's how you get a literature that's structurally tilted before a single editor ever makes a bad call.

Iris Holm: But here's what that architecture actually means — Rosenthal wasn't just naming a metaphor. He was saying: the bias isn't at the gate. It's way earlier.

Cyrus Reed: Right, and 1979 is when he formally put it down — the file drawer problem. Which is — wait, it's not just a name, he actually quantified it. Like, how many unseen null results would you need sitting in file drawers to cancel out a given body of significant findings? He ran that math.

Iris Holm: How many is it?

Cyrus Reed: Depends on the literature, but the point is — it's almost always achievable. The number of null results that would need to exist to wipe out most celebrated findings is not astronomically large. Which means the comfortable assumption that 'published equals established' is just... wrong.

Iris Holm: And the 60-point gap is the actual measure of that pre-submission suppression.

Cyrus Reed: Sixty points. Not in rejection — in who even bothers to write it up. And there's actually — okay, picture a postdoc, conference poster half-finished, she's got results from a replication study, null across the board. Her advisor's not telling her to bury it. Nobody is. She just — she looks at her CV, she counts the publications she needs before her contract ends, and she opens a new file for the next experiment instead. That's the file drawer problem. It's a career calculation, not a deception.

Iris Holm: So the bias is coming from researchers. Not at editors.

Cyrus Reed: Which is — I mean, that's the uncomfortable inversion, right? Because every reform conversation starts with journals. Change what editors accept. But positive results are only forty points more likely to be published once submitted. The sixty-point gap is pre-submission. The editors never even see most of the null results to reject them.

Iris Holm: Sterling's 1995 audit confirmed the same pattern as 1959. Thirty-six years. No movement.

Cyrus Reed: Nothing changed! Sterling, Rosenbaum, and Weinkam — same finding. Which means every intervention in between did essentially nothing to the shape of the literature. And the reason is — actually, no, wait — the reason isn't that the interventions were bad. It's that they mostly targeted journals, and the suppression was already done before the journal saw anything.

Iris Holm: So reforming editorial policy is treating a downstream symptom.

Cyrus Reed: Exactly — and there's a structural layer underneath this that makes it even harder to escape, something about why individual researchers literally cannot opt out even when they want to, the publish-or-perish machinery and how journal impact factor feeds tenure decisions. That's — we should hold that thread, because it changes what fixing this even looks like.

Iris Holm: So the question is: what would it actually take to fix the pre-submission problem, if the journal isn't the chokepoint?

Cyrus Reed: That's the trap — because opting out isn't actually available. Like, imagine a junior faculty member, third year, grant renewal coming up. She could decide to submit her null result. She could. But her tenure file goes before a committee that counts publications in high-impact journals, and 'impact factor' is just — wait, it's average citations per article, right, which means the journal is already optimizing for novelty because novel papers get cited more, which raises the number, which attracts bigger names, which raises it again. She's not fighting one gatekeeper. She's fighting a loop.

Iris Holm: Three independent nodes. Publish or perish, journal impact factor, and funders who reward track records of high-profile publications — which are themselves products of the same filter.

Cyrus Reed: And each one applies the same novelty-and-significance premium independently. So even if she convinces her department chair, the grant agency is still running the old calculus.

Iris Holm: Can she just... decide to do it differently?

Cyrus Reed: Not without real career cost. That's — no, that's the honest answer. The Iestyn Williams stakeholder interviews document exactly this: funders, publishers, and researchers all separately describing the same squeeze, from their own angle, in health services research. It's not one villain. Everyone's responding rationally to the system they're actually in.

Iris Holm: Which is what makes Ioannidis so uncomfortable. PLOS Medicine, 2005 — he's not accusing anyone. He's running the math.

Cyrus Reed: Right — and the math says: given realistic prior probability that a hypothesis is true, realistic statistical power, and the bias we've been describing, most published findings are probably false. Not 'sometimes wrong.' Most. That's a structural prediction, not a scandal.

Iris Holm: Wait — most? That's the actual claim?

Cyrus Reed: Most. Under realistic conditions. And Ioannidis isn't saying every paper is fabricated — he's saying the architecture we just described, the incentive misalignment where rewards track novelty and impact factor and not replicability or effect size, makes false positives the statistically expected output. The system is doing what it was built to do.

Iris Holm: So the asymmetry is load-bearing. It costs her nothing to not submit the null result. She loses nothing she already had. Submitting it and being seen as unproductive — that costs something real.

Cyrus Reed: And that gap — between the cost of silence and the cost of honesty — that's actually what Sterling's 1959-to-1995 stability is measuring. Forty years of that same rational choice, compounding.

Iris Holm: So reform that doesn't touch what's rewarded at tenure time just slides back to the same equilibrium.

Cyrus Reed: And that's exactly why Registered Reports feel like — wait, okay, they are a real thing. Peer review happens before data collection. The journal commits to publishing based on the design, not the outcome. So a null result literally cannot get you rejected. And the data backs it up — studies of Registered Reports show higher rates of null-result publications and less outcome-switching compared to standard tracks. That's not nothing.

Iris Holm: It's not nothing. But it fixes analytical flexibility. It doesn't fix novelty-as-reward.

Cyrus Reed: Say more — because I want to believe pre-registration actually solves it, but I think you're pointing at something I can't quite — okay, like, even if the null result gets published, does anyone cite it? Does it move the tenure file?

Iris Holm: No. That's the gap. A Registered Report can print your null finding. But journal impact factor is still average citations per article. Novel positive results get cited. They pull the number up. The Registered Report with its careful null sits there. Adoption stays limited because the career math hasn't moved.

Cyrus Reed: The plumbing changed and the pressure didn't. The bias isn't editors. It's a researcher on a Tuesday morning deciding the null result isn't worth the fight. That's the Philosophical Transactions of the Royal Society in the seventeenth century forward: the system has been selecting for signal the entire time. Peer review is old enough that no living scientist built it.

Iris Holm: And knowing that doesn't make it easy to fix. It just tells you where the fix has to land — not the plumbing. The incentives.

Cyrus Reed: Tuesday morning again. Full circle. She's still sitting there with her null result, and the Registered Reports format exists now, pre-registration exists, and she still has to decide if it's worth it given her contract. Nothing we named today actually changes that morning for her.

Iris Holm: That's probably the most honest place we could end up. Unresolved but — yeah. Not a bad morning's work.

How publication incentives systematically bias toward surprising results over robust ones · Onpode