Onpode
Cover art for Why peer review catches obvious errors but misses subtle methodological flaws

Why peer review catches obvious errors but misses subtle methodological flaws

August 2, 2026 · 15 min

Hugo Vance & Lila Soto

In BMJ embedded-error studies, some peer reviewers caught zero of eight or nine planted mistakes, and only 10% caught four or more. Peer review functions as a gatekeeping floor — filtering gross errors — not a quality guarantee. Three flaws it structurally cannot detect: post-hoc sample size justification, selective outcome reporting, and p-hacking.

Peer review is the gatekeeping mechanism through which manuscripts are evaluated by independent field experts before publication. Originating in the 17th century alongside the professionalization of science — with early formalization through journals like the Philosophical Transactions of the Royal Society — it has become the dominant credentialing process in academic publishing.

0:0014:39
Get the next episode on Science

Follow it free — new episodes land in your feed.

Or make your own — any topic, in minutes

More Onpode episodes on Science

About this episode

Peer review has a reputation problem it didn't entirely earn — and a structural problem it hasn't been honest about. This episode starts with a number: zero. In the BMJ's embedded-error studies, some reviewers caught none of the deliberate mistakes planted in manuscripts. Only ten percent caught four or more. The episode uses that result not as an indictment of individual reviewers, but as a lens on what the system was actually designed to do — and how far the public understanding has drifted from that design. The conversation works through three specific failure modes that are invisible to any reviewer holding only the final manuscript: post-hoc sample size justification, selective outcome reporting, and p-hacking. In each case, the problem isn't that the reviewer missed something — it's that the relevant information was never in front of them. From there, the episode examines the full menu of proposed reforms and finds a consistent pattern: every fix trades one failure mode for another, and none of them touches the underlying problem, which is the label itself. 'Peer-reviewed' is one undifferentiated signal. A rapid conference review and a six-month double-blind trial both produce it. The mechanism can change entirely; the signal the public receives stays the same. That's the thread the episode follows to its genuinely unresolved end.

Frequently asked

How effective is peer review at catching errors in research papers?

Peer review catches obvious errors but misses subtle methodological flaws. BMJ embedded-error studies found some reviewers caught zero of eight or nine planted mistakes; only 10% caught four or more. Reviewers average 2.7 manuscripts per submission, unpaid, and see only the final version — not the original study protocol.

What types of research fraud can peer review not detect?

Peer review cannot detect post-hoc sample size justification, selective outcome reporting, or p-hacking, because all three leave no visible artifact in the final manuscript. A reviewer holding only the submitted paper sees a coherent, internally consistent study — the discarded protocols and null findings were never submitted for review.

Does double-blind peer review eliminate bias?

Double-blind peer review reduces bias substantially — 74 to 90 percent of reviews contain no correct guess of author identity — but the fix breaks down in small, specialized fields where reviewers can identify a lab from its methodology alone, making double-blind review functionally ineffective in certain subfields of neuroscience and microbiology.

What does 'peer-reviewed' actually guarantee about a published study?

The label 'peer-reviewed' guarantees only that a manuscript cleared a minimum-coherence gatekeeping process — it does not certify depth, rigor, or replication. A 45-minute rapid conference review and a six-month rigorous double-blind trial both produce the same four-word label, with no public disclosure of which process was actually applied.

What reforms exist for peer review, and do they fix the core problem?

Proposed peer review reforms include registered reports, data-sharing mandates, open post-publication review, distributed review (ALMA used over 1,000 reviewers for 1,497 proposals at Cycle 8), and ORCID reviewer recognition. Each trades one failure mode for another, and none changes the undifferentiated public signal produced by the 'peer-reviewed' label.

Grounded in 12 sources
Rejoinder: The ICML 2023 Ranking Experiment: Examining Author Self-Assessment in ML/AI Peer Review · arxiv.org
The More You Automate, the Less You See: Hidden Pitfalls of AI Scientist Systems · arxiv.org
Analysis of the ALMA Cycle 8 Distributed Peer Review Process · arxiv.org
Effectiveness of Anonymization in Double-Blind Review · arxiv.org
Does the use of open, non-anonymous peer review in scholarly publishing introduce bias? Evidence from the F1000 post-publication open peer review publishing model · arxiv.org
Optimizing Peer Grading: A Systematic Literature Review of Reviewer Assignment Strategies and Quantity of Reviewers · arxiv.org
The Extent and Consequences of P-Hacking in Science | PLOS Biology · journals.plos.org
Does the disconnect between the peer-reviewed label and reality explain the peer review crisis, and can open peer review or preprints resolve it? A narrative review | Naunyn-Schmiedeberg's Archives of · link.springer.com
Can I trust this paper? | Psychonomic Bulletin & Review | Springer Nature Link · link.springer.com
Statistical reviewing for a medical journal: summary of workshop proceedings | Research Integrity and Peer Review | Springer Nature Link · link.springer.com
Peer Review in Scientific Publications: Benefits, Critiques, & A Survival Guide - PMC · pmc.ncbi.nlm.nih.gov
Opening the black box: Understanding the editorial and peer review journey of a manuscript - PMC · pmc.ncbi.nlm.nih.gov
Read transcript

Hugo Vance: You have been unusually quiet on the messages this week, which usually means you found something that bothers you.

Lila Soto: Ha — yeah, I mean, I found a number and I cannot make it mean what I thought it would mean. We're doing peer review today, and the number is zero.

Hugo Vance: Zero.

Lila Soto: In the BMJ's embedded-error studies — British Medical Journal, they planted deliberate mistakes in manuscripts — some reviewers caught zero of eight or nine errors. And only ten percent caught four or more. I kept waiting for that to feel like a bug in a specific implementation. It doesn't. It feels like the whole point was never what we said it was.

Hugo Vance: I'd push back on that framing immediately. Because the Philosophical Transactions of the Royal Society — 17th century, the origin point — was built to accelerate trust, not to guarantee truth. Those are not the same project.

Lila Soto: Okay but who told the public that?

Hugo Vance: Nobody. And that is the crux of it.

Lila Soto: So peer review is — and I want a plain sentence for this — peer review is a gatekeeping floor, not a quality guarantee. That's the thing we've gotten catastrophically wrong in how we talk about it.

Hugo Vance: Yes. Think of it this way: a smoke detector is not a fire suppression system. It signals that something may be burning. Peer review was always the alarm. We decided, somewhere along the way, to call it the sprinkler.

Lila Soto: Huh — I like that. But a smoke detector with dead batteries is a different problem than a smoke detector that was never designed to suppress fire.

Hugo Vance: Now that is the interesting distinction. You see, the dead-battery problem — that is the 2.7 reviews per manuscript, unpaid, on top of full academic workloads. That is a structural failure, not a conceptual one.

Lila Soto: I'm not sure I can separate those. If the batteries are always dead, at what point is 'smoke detector' just a name we gave a piece of plastic on the ceiling?

Hugo Vance: That piece of plastic on the ceiling — yes, that is where I'd be cautious, because it collapses two distinct failures. The concept was sound. What we've broken is something else entirely. There are three specific ways a manuscript can be fraudulent on arrival and appear, to any reviewer, completely clean.

Lila Soto: Clean how? Like, technically coherent?

Hugo Vance: Picture a cardiologist. Thursday evening. Two glasses of wine, forty minutes, manuscript on heart failure outcomes. The power calculation looks adequate. The methods section reads logically. She cannot know — she has no instrument that could tell her — that the authors ran that study for eighteen months, hit an underpowered wall, and then retrofitted the sample size justification after the data were in. That is post-hoc sample size justification. It is statistically defensible on paper. The original protocol is not in front of her.

Lila Soto: Wait — the protocol is just... not submitted?

Hugo Vance: Not routinely, no. And that is before you get to selective outcome reporting — where the null findings simply do not appear in the manuscript. The paper is internally coherent. Nothing contradicts itself. The reviewer sees a complete story, because the incomplete story was never handed to them.

Lila Soto: So the indictment isn't that the cardiologist failed. It's that we built a system and then withheld half the information from the person we asked to validate it. If we know reviewers can't see behind the manuscript — and we've known this — why hasn't the question become what we hand them?

Hugo Vance: Well, COPE — the Committee on Publication Ethics — has standards for exactly this. Expected reviewer conduct, disclosure obligations. I think it is genuinely well-intentioned. But you cannot write an ethical guideline that resolves a structural labor economics problem. A reviewer completing 2.7 manuscripts unpaid is not going to chase a protocol that was never attached.

Lila Soto: Right — but COPE setting standards for conduct and COPE fixing the incentive structure are completely different projects, and I'm not sure we've been honest about which one they're actually doing.

Hugo Vance: Quite.

Lila Soto: And then there's p-hacking, which is — I mean, the same structural feature, right? Researchers test a dozen outcomes, report the two that hit significance, and the reviewer sees one clean manuscript with two significant findings. There is no artifact of the fishing expedition visible anywhere in what gets submitted.

Hugo Vance: All three of them — post-hoc justification, selective outcome reporting, p-hacking — they share that single feature. Invisible to anyone holding only the final version. You see, the BMJ result — two of eight or nine errors caught on average — that is not evidence of reviewer incompetence. It is evidence we assigned them a job they were structurally never equipped to do.

Lila Soto: I don't disagree with that framing — but it almost lets us off the hook too easily. If the job was always impossible, someone decided to call it sufficient anyway. That's not a design flaw. That's a choice that has consequences every time the words 'peer-reviewed' appear on a vaccine study or a climate paper and someone decides that phrase is enough.

Hugo Vance: That choice — and it was a choice — connects directly to what the reforms are actually offering. Because the moment you start down the list of fixes, every single one trades one failure mode for another.

Lila Soto: That's the part I can't get past. Like, ALMA ran distributed peer review at Cycle 8 — over a thousand reviewers, 1,497 proposals, nearly fifteen thousand ranks and comments. Largest deployment in astronomy. And my first reaction was: that's incredible. And then I thought — wait, does spreading the load actually fix depth, or does it just make the shallowness invisible because you're averaging across a thousand people?

Hugo Vance: The averaging problem. Yes. That is the correct instinct.

Lila Soto: And then double-blind review — I mean, the numbers are genuinely good. 74 to 90 percent of reviews contain no correct guess of author identity. That's real bias reduction. But it collapses in small fields where everyone knows everyone anyway, so the fix only works where the problem is already less acute?

Hugo Vance: In certain branches of neuroscience, certain subfields of microbiology — double-blind is functionally theatre. Everyone can identify the lab from the methodology alone.

Lila Soto: And F1000Research goes the opposite direction — open, post-publication, reviewer names attached. Transparent. And what happens? Sequential-influence bias. National-affiliation bias. The first reviewer's comment shapes every reviewer after them. So transparency introduces its own distortion.

Hugo Vance: And AI-assisted review adds pattern-matching errors that neither editors nor authors are yet equipped to audit. No one in the chain has the tools to catch what the AI gets wrong.

Lila Soto: So — every reform. Every single one. ORCID recognition at least names the labor, gives reviewers credit for the work, but it doesn't change that the work is still unpaid and still structurally capped at 2.7 manuscripts per submission.

Hugo Vance: I'd call that a portfolio, not a paradox. No single mechanism was ever going to be sufficient. The question is whether layering — registered reports, data-sharing mandates on top of preserved gatekeeping — whether that stack is adequate.

Lila Soto: No, I don't buy that.

Lila Soto: Because layering fixes the mechanism. It does nothing to the label. 'Peer-reviewed' still lands on the vaccine headline, the climate paper, and the public reads that phrase and stops. The stack underneath could be completely different and the signal they receive is identical.

Hugo Vance: Well — I think that is where I have to concede something uncomfortable. The label has become culturally decoupled from whatever the mechanism actually is. You could reform the process entirely and the public-trust problem would survive untouched. That part, I'd say, is the darker thread — and we're not quite at it yet.

Lila Soto: The expertise that would actually fix shallow review is exactly what makes qualified reviewers overloaded. The more specialized the field, the fewer people who can do it — and the more each of those people is already drowning.

Hugo Vance: You see, that is not a paradox you solve by adding reviewers. You solve it by changing what you ask any single reviewer to certify — and we have not done that.

Lila Soto: And what we haven't named is what that looks like from the outside — not the mechanism failing, but the label surviving the failure completely intact. ICML 2023 ran an experiment on this, sort of. They looked at calibration between how authors ranked their own work and how reviewers ranked it in AI and machine learning peer review. And the gap was — I mean, it was poor. Even in a high-rigor field. The author thinks it's a strong submission, the reviewer disagrees, and neither of them is necessarily wrong — they're just operating on completely different frameworks.

Hugo Vance: Yes. And the paper that cleared that process carries the same label as one that cleared exhaustive double-blind scrutiny over six months.

Lila Soto: The same four words.

Hugo Vance: I want to be careful here, because — well, this is the thing I have to concede. I have argued, consistently, that the fix is layering. Registered reports, data-sharing mandates, post-publication scrutiny on top of preserved gatekeeping. And I hold that. But there is a narrower problem I have not adequately answered, which is this: you can reform the entire stack and the public-trust problem survives untouched. Because the label is one undifferentiated signal. A rapid conference review and a rigorous double-blind trial both produce it. The mechanism changed. The signal did not.

Lila Soto: That's — yeah. That's the concession I've been waiting for.

Hugo Vance: But I am not saying the fix is to blow up peer review. The gatekeeping function — filtering gross errors, enforcing minimum coherence — that is a necessary floor. Abolishing it is not the answer. The answer is to be honest, publicly, about what the label does and does not certify. Which we have never done.

Lila Soto: I don't disagree with that. But I'm wondering whether 'being honest about the label' is actually achievable without — I mean, who does that? Who tells the person reading the climate headline that this particular peer review was a 45-minute skim and that one was six months of scrutiny? There's no infrastructure for that distinction.

Hugo Vance: No. And that is the part I cannot fully resolve — I can see the shape of what would need to exist, a tiered disclosure, some version of review-depth transparency attached to the label itself. But the credential function has calcified. Journals, institutions, tenure committees — they have all built on that single signal.

Lila Soto: Which means mechanism reform alone is insufficient by definition. If the label is the problem, fixing what happens underneath it doesn't touch what the public receives.

Hugo Vance: That I will not contest.

Lila Soto: I'm not trying to win the argument — I just think it means the stack you're describing, registered reports, data-sharing, preserved gatekeeping — that's producing better science. It may not be producing better-informed readers of science. Those are two different outcomes.

Hugo Vance: Yes. And I think that is where we actually land — peer review remains a necessary floor, not a sufficient ceiling. The goal is to supplement it, not to replace it. But the credentialing function and the quality-control function have come apart, and nobody in the chain — not COPE, not journals, not the ICML program committee — has been willing to say that plainly to anyone outside the field.

Lila Soto: Which brings me back to the number I started with. Zero. Some reviewers caught zero of eight or nine errors. And I thought that was the indictment when I found it. Now I think — I mean, the zero almost isn't the point. The point is that the person who caught zero and the person who caught four both produced the same output. The same label.

Hugo Vance: That is the quiet version of the problem. Yes.

Lila Soto: And if the next generation of researchers just... skips it — posts preprints, waits for rapid community feedback, builds reputation on replication records instead of journal placement — the credentialing system doesn't get reformed. It just gets bypassed. And I'm genuinely not sure whether that's a disaster or a correction.

Hugo Vance: Well — I think it is both, depending on what you lose in the transition. The floor disappears with the institution. Preprint communities do not have a COPE. They do not have a shared standard for what minimum coherence even requires. You get speed and you get distribution, and you lose the thing that was — however imperfectly — enforcing a basement.

Lila Soto: A floor nobody can see through.

Hugo Vance: I won't walk away from it, no. Imperfect floors are still floors.

Lila Soto: I know. And I don't think I'm arguing against it — I'm just not sure defending it looks different from the outside than pretending it works. That's the part that stays open for me.

Hugo Vance: It stays open for me too. Which is not something I say often about an institution I've taught for thirty years. That's — well. That's probably enough for one Tuesday.