June Hadley: Can I just say, before anything else — I've been thinking about this topic as a thought experiment in invisible architecture, and I think that framing is going to be useful for us today.
Cleo Rios: Invisible architecture — yes, because that's what it is, and nobody talks about it that way. How was your week, by the way, are you — actually, no, I cannot wait, I need to say the thing.
June Hadley: Say the thing.
Cleo Rios: Peer review is the process that decides what counts as knowledge. Not what's true — what gets to be called true, officially, in the record. That's the gate we're talking about. And I'm obsessed with it because it's the most powerful invisible institution in science and most people have no idea it exists.
June Hadley: And here's the piece that I think crystallizes why it matters — and why it's stranger than it looks. The International Conference on Learning Representations, ICLR, switched from single-blind to double-blind peer review in 2018. Researchers then went through 5,027 submitted papers from that natural experiment. And the finding was that double-blind review significantly reduced scores for high-prestige authors. The prestige bias, the tendency to favor famous labs, shrank in the actual scores.
Cleo Rios: Right — but acceptance rates were unchanged. Which is the part that made me put my phone face-down on the table for a minute.
June Hadley: It's — I mean, it genuinely should have been reassuring. The anonymity model worked. Single-blind versus double-blind, that's a real difference in who knows what. And the scores reflected it. So why didn't the outcomes?
Cleo Rios: Because the bias is not living in the scores. That's the whole thing. The gatekeeping function — the part that determines what counts as validated knowledge — that part is operating somewhere the anonymity couldn't reach.
June Hadley: Which is a very different problem than 'reviewers are sometimes unfair.'
Cleo Rios: Completely different problem. And a much harder one.
June Hadley: So the question sitting underneath all of this — if reforming the scoring doesn't move the gate, are we reforming the wrong thing? And does that mean the mechanism is broken, or does it mean it's working exactly as designed?
Cleo Rios: Oh, that second option is the one that keeps me up. Because 'working exactly as designed' is a much scarier answer.
June Hadley: And 'working exactly as designed' — I think that's actually where we need to slow down, because it changes what we're even looking at. Let me try the plain-language version first. Peer review is like having three experienced cooks taste your dish before it goes on the menu. They can say it needs work, or it's not ready. The idea is: nothing reaches the public without that tasting. That's the whole intuition.
Cleo Rios: Right — and that works beautifully, until you realize the cooks already knew whose dish it was before they tasted it. That's the ICLR thing. They made the tasting blind — genuinely blind, author names hidden — and the scores changed. High-prestige authors got knocked down. The cooks rated the dish differently when they didn't know the chef.
June Hadley: But the dish still made the menu.
Cleo Rios: At the same rate. Which means — the score was never actually the gate. The score was, like, the performance of the gate. Something else was deciding.
June Hadley: So where is that something else? I mean, if the 5,027 papers they studied showed unaffected acceptance rates even after bias measurably shrank in the scores — what are we even pointing at when we say 'the review decided'?
Cleo Rios: Okay — upstream. It has to be upstream. Who got invited to submit in the first place? Whose question was legible as a serious research question before a single reviewer opened a file? That's not in the score. That's not touchable by anonymity at all.
June Hadley: Which is — wait, that's actually the more unsettling version of prestige bias. Because anonymizing the review process assumes the bias enters when the reviewer reads the name. But if it enters when 'significance' itself gets defined — whose work gets framed as worth reviewing — then double-blind is solving for step four of a ten-step problem.
Cleo Rios: Step four at best.
June Hadley: Now, does that mean the reform was useless? I don't think I'd say that — actually, no, I'd say something more uncomfortable. The reform revealed exactly where the bias wasn't, which tells us something real. It's just not the thing anyone wanted to hear.
Cleo Rios: It's like — you fixed the lock and then discovered the wall has a door-shaped hole in it. The lock works perfectly now! Great. Irrelevant.
June Hadley: And the door-shaped hole is structural. It's in how fields decide what counts as a serious question, what counts as a legitimate method, whose framework is the default framework. That's not a flaw you patch with an anonymity model.
Cleo Rios: No, because the conservatism isn't accidental — that's the mechanism doing its job, which is protecting the current shape of knowledge. And if that's true, then asking peer review to also be the engine of genuinely novel work is asking it to be two opposite things at once.
June Hadley: And that structural thing — that's what makes me want to go one layer further, because there's an empirical record here that I don't think we've actually sat with yet. The conservatism question is real, but underneath it there's a more basic problem. Two reviewers reading the exact same manuscript frequently reach opposite conclusions. Not edge cases — frequently.
Cleo Rios: Wait, like — disagree on accept versus reject?
June Hadley: Accept, reject, major revision — the inter-reviewer reliability is genuinely low. And on top of that, major methodological errors get missed. So the validity problem isn't just 'reviewers have opinions.' It's that reviewers are not reliably catching bad science.
Cleo Rios: No — stop — if that's true, then what is the gate actually filtering? Like, if the cooks are tasting the same dish and one says rancid and one says perfect, and also neither of them noticed it had raw chicken in it — the gate is not doing the thing we said the gate does.
June Hadley: That's the validity question exactly. And it's — I mean, it sits right next to the reliability problem but it's distinct. Low reliability means reviewers disagree. Low validity means they agree on the wrong thing, or miss the right thing entirely.
Cleo Rios: Both at once is kind of a disaster.
June Hadley: And then the NLP studies come in on top of that, which is actually where I think the story gets stranger. Because researchers ran sentiment analysis on large corpora of peer review text — the actual prose reviewers write — and the tone varies. Systematically. With author gender, race, institutional affiliation. Not the score. The language.
Cleo Rios: Okay, paint me the scenario because I need to feel this one. Give me a person.
June Hadley: Researcher at a regional state school. She's working at the intersection of anthropology and machine learning — which doesn't map cleanly onto the existing boxes at a top venue. Formats everything perfectly. Methodology is solid. Submits. Three weeks later, vague rejection. And one review reads measurably colder than the others — not a different score, just colder prose — without naming a single methodological flaw.
Cleo Rios: That review is doing something. It's not catching bad science — it's not even attempting to. It's performing skepticism through tone, and that tone is carrying the cost. And anonymity — the reviewer is anonymous, she's not — anonymity cannot touch what lives in the prose itself.
June Hadley: Right — because demographic bias in review tone operates before methodology is assessed. It's in the framing, the warmth, the way a question gets characterized as interesting versus peripheral. That's not fixable by hiding names. The bias is already working.
Cleo Rios: And honestly, that's the part that makes what's coming next in this conversation feel even heavier — because when you layer in who peer review has historically excluded by design, not accident, the picture gets a lot darker than 'low reliability.'
June Hadley: And darker is — I think that's the right word, actually. Because Ncube, 2025, names this explicitly. The argument isn't just that peer review is biased. It's that methodological gatekeeping operates as epistemic marginalisation — that disciplinary conventions encode dominant-culture assumptions as if they were universal scientific standards. That's a different claim.
Cleo Rios: Epistemic marginalisation — meaning the method itself is the message. Like, 'your findings might be correct but your framework is not legible to us, so you don't exist in the record.'
June Hadley: Which is not a reviewer being mean. It's the structure deciding whose questions count as scientific questions.
Cleo Rios: And the structural move everyone avoids is this — it's not who signs the review. It's who gets invited to review. That's a composition problem. The panel itself encodes the bias before a single paper lands in anyone's inbox.
June Hadley: And anonymity — double-blind, open, any model — leaves that entirely intact.
Cleo Rios: Completely untouched. Now here's what makes this vertigo-inducing to me — the proposed fix right now, the thing getting serious traction, is integrating Generative AI into the review process. And I'm sorry, but — an AI trained on existing literature would be the most conservative reviewer imaginable.
June Hadley: Wait — say that again, because I want to make sure we land this precisely.
Cleo Rios: Biele and Kowalski, 2026 — they argue that integrating Generative AI into peer review would systematically flag paradigm-shifting work as noise. Because the AI learned from what already got published. So genuinely novel findings — the anomalies that break frameworks — would read to the model as error.
June Hadley: So the conservative bias peer review already has — it would be, I mean — it would be formalized. Baked into the algorithm.
Cleo Rios: And Biele and Kowalski actually name four things they say only human reviewers can do — verifying ground truth, arbitrating ethics, communicating uncertainty, and curating the paradigm-shifting anomalies the AI would reject. Those four things are the whole ballgame.
June Hadley: The fourth one especially. Curating the anomaly — that requires knowing it's an anomaly worth keeping, not noise worth discarding. An AI can't do that without the conceptual history to recognize what breaks the frame.
Cleo Rios: But — and I have to hold the other side of this or we're being unfair — the system without peer review also fails. Badly.
June Hadley: The telehealth case.
Cleo Rios: A report on virtual physical therapy. Disseminated without peer review. Flawed methodology, spin, unsupported claims — and it demonstrably influenced clinical practice and healthcare policy. Real patients. Real decisions. Made on research that had never been through the gate.
June Hadley: So it's not a broken system versus a good one. It's two different failure modes. Gatekeeping bias on one side, unfiltered methodological error reaching policy on the other.
Cleo Rios: And F1000Research is trying to thread that needle — open, post-publication, non-anonymous review. Which is a real live experiment in whether transparency trades one set of problems for another.
June Hadley: The tradeoff is just genuinely hard. Preprint dissemination, the explosion of online-only journals publishing with little or no review — that's faster access, yes. But the error-correction function disappears with it.
Cleo Rios: Which means we're not trying to fix a broken mechanism — we're trying to choose which way we're willing to fail. And I don't think anyone's said that plainly enough yet.
June Hadley: And that's — I mean, that's the thought I can't settle. Because if we're choosing which way we're willing to fail, the question underneath is: can any single mechanism simultaneously protect rigor, enable innovation, suppress bias, and hold public trust? Those aren't four versions of the same job. They're — actually, some of them are structurally opposed.
Cleo Rios: We keep reforming the review. We never reform the reviewer pool. That's the move nobody makes.
June Hadley: No, and that's — that's where the ICLR finding lands differently now than it did at the top of this conversation. The scores changed. The outcomes didn't. And we know now that the bias was upstream — in language, in whose question was legible, in who was in the room doing the reviewing at all. So the 'seal of quality' that underpins public trust in science, the thing the American Institute of Biological Sciences points to, the thing the Higher Learning Commission treats as a standard — that seal is being issued by a process we've just spent an hour demonstrating cannot do all of what we've asked it to do.
Cleo Rios: I don't know what the answer is. I genuinely don't. And I think that's the honest place this lands.