Lila Soto: Okay, I have to tell you the thing that's been nagging me since I went down this rabbit hole — there's a researcher, doing careful, rigorous work, replicates a famous study, finds it doesn't hold. Solid methodology. And then the paper goes nowhere. Meanwhile the original finding has, like, five hundred citations.
Lila Soto: Right — and a grant, and a tenure case. And what I keep wondering is whether that's a flaw in the system or just... the system working exactly as designed.
Iris Holm: It's the system working as designed. That's the uncomfortable part.
Lila Soto: Which is what today is actually about — this incentive misalignment between what rewards individual scientists and what would actually benefit science as a whole. And the place where that becomes concrete is the Reproducibility Project: Psychology. Brian Nosek, the Center for Open Science, 97 studies — and only 36% replicated.
Iris Holm: Thirty-six. Not thirty-six percent of a fringe. Ninety-seven high-profile studies.
Lila Soto: Hm, and what gets me is — this isn't even a secret. Surveys of more than 1,500 scientists found over 70% had already personally failed to reproduce someone else's results. It's the normal experience. And we still treat replication work like it's career poison.
Iris Holm: Because it is. That's not rhetoric — the publish-or-perish norm structurally punishes the one activity that would actually stress-test the literature.
Lila Soto: Yeah, so the question we're really trying to work out — why does a system built by people trying to do good science end up selecting against reliable science?
Iris Holm: Good faith, bad architecture.
Lila Soto: Good faith, bad architecture — yeah, that's exactly the frame.
Iris Holm: Now: what does it take to change the architecture when every institution that would have to change it is also winning under the current rules?
Lila Soto: But that's the part I keep snagging on — like, how did the architecture get wired that way in the first place? Because someone had to decide that Journal Impact Factor meant something.
Iris Holm: Think of it like a movie review site where the score only counts how many people clicked the review. Not whether the movie was good. Just clicks. Now imagine studios started greenlighting movies purely to get clicked.
Lila Soto: Oh. That's — yeah, that's exactly it.
Iris Holm: Impact Factor counts citation frequency. Not accuracy. And then universities took that number and built tenure decisions on top of it. Hiring decisions. Grant allocations. The metric measures buzz, and we made buzz the definition of quality.
Lila Soto: And it doesn't stop at the journal level — I mean, the same logic runs through h-index, citation count. All the bibliometric indicators are doing the same work at the individual researcher level. So it's kind of... stacked.
Iris Holm: Stacked is the right word. And here's the mechanism nobody says plainly: a replication study that overturns a high-profile finding is, by construction, less citable than the original. Because the original has five years of downstream papers all citing it.
Lila Soto: Wait — so the corrective is structurally worth less than the error it corrects?
Iris Holm: Always. The original gets cited by everyone who built on it. The correction gets cited by... people writing about corrections. PNAS published research making exactly this point — framing the whole thing as systemic, not moral. It's not that scientists are lying. The market design produces this outcome from honest behavior.
Lila Soto: Which is almost worse? Like, if it were just bad actors you could fix it with enforcement. But if it's rational behavior inside a broken structure — yeah, I don't know what the lever is.
Iris Holm: The lever is what actually determines whether someone gets hired. And right now that's Impact Factor and citation count. Preregistration doesn't touch that. Neither does open peer review.
Lila Soto: Hm. So top journals preferentially publish the highly novel study — the one that's never been done before — even if that novelty makes it harder to verify. And that study lands the citations, which feeds the h-index, which feeds the tenure case. The whole chain runs on novelty.
Iris Holm: And the researcher who does careful replication work — solid methodology, null result — gets rejected everywhere, struggles to fund the next project. The question is whether any institution is actually willing to change what it rewards at the moment that counts. The hiring meeting. The tenure vote.
Lila Soto: And that's where it stops being abstract — because she's not sitting there plotting. She's just doing the math on her own survival. The postdoc, eight months left on her contract, null result in front of her. What does the system train her to do next?
Iris Holm: Run the subgroups.
Lila Soto: Run the subgroups. Try a different control variable. Slice the data a different way until something clears p less than 0.05. And that's p-hacking — not fabrication, just... iteration until the number cooperates.
Iris Holm: And individually? Totally rational. Publication bias means null results don't get published. No publications, no renewal. The math is clear.
Lila Soto: I mean — it's not even, like, a moral failure. The structure produced this outcome. She didn't invent the incentive.
Iris Holm: Right. And then the companion move — once the data shows something unexpected, you write the introduction as though you predicted it. That's HARKing. Hypothesizing After the Results are Known. Exploratory finding, dressed as confirmatory.
Lila Soto: Wait — so the paper reads like a hypothesis test, but it's actually pattern-matching on the same dataset that generated the pattern.
Iris Holm: Every time. And publication bias closes the loop — because the journals are only seeing the sliced version, the one that cleared 0.05. The null result stays in a drawer. So the literature fills up with findings that look solid and aren't.
Lila Soto: Which is actually — okay, I want to say this carefully — that's a market-design failure, not a misconduct problem. The reward structure made this individually rational and collectively self-destructive. Like, she did what the system told her to do, and the system got worse.
Iris Holm: If you enforce against individuals without changing the reward structure, you just get more careful p-hacking. The behavior is load-bearing. It's holding the career up.
Lila Soto: Hm. And what nobody talks about enough — she probably knows. Like, she's not confused about what she's doing. She's made a calculation.
Iris Holm: Which is the bleaker version of the problem.
Lila Soto: Yeah. And the numbers that actually show how deep this goes — the Reproducibility Project, what the replication rates look like across fields, what preregistration and Registered Reports can actually fix and what they can't — that's kind of where this gets even harder to sit with.
Iris Holm: But here's where the numbers get weird — because if p-hacking is rational everywhere, the replication rates shouldn't look this different across fields. Psychology: 36%. Economics and social science: 60 to 70. Same broken incentive structure. Very different outcomes.
Lila Soto: Wait — that's a huge gap. Why does economics look so much cleaner?
Iris Holm: Nobody's settled it. But my read — economists use larger datasets almost by default, administrative records, census data. Harder to slice into significance when n is a hundred thousand. Psychology runs a lot of sixty-person lab studies where the wiggle room is enormous.
Lila Soto: So it's not that economists are more virtuous — it's that the data structure makes p-hacking harder to execute cleanly.
Iris Holm: Which changes the whole diagnosis. If this were purely a culture problem, you'd fix it with culture. But if method architecture drives the number — then the fix has to be structural too.
Lila Soto: And then there's preclinical cancer biology sitting at around 40%, which is — I mean, that's not a methodology seminar problem. That's researchers building drug trials on foundations that don't hold half the time.
Iris Holm: 40% in preclinical work, yes. And that's where the abstraction ends.
Lila Soto: Like — a compound gets flagged as promising because the lab result looked solid. It enters trials. The lab result was one of the 60% that didn't hold. And you're asking patients to be in that trial.
Iris Holm: Frankly, that's the case for why the reform tools matter. Preregistration — you register your hypothesis, your design, your analysis plan on the Open Science Framework before you touch the data. The p-hacking opportunity mostly closes.
Lila Soto: And Registered Reports go further, right? The journal accepts or rejects based on methodology before results exist. In-Principle Acceptance. So the outcome literally cannot determine whether it gets published.
Iris Holm: That's the mechanism. You decouple the publication decision from the result. Which is — look, that's the only intervention that actually attacks publication bias at the journal level, not just the researcher level.
Lila Soto: Hm. But they're voluntary. The Open Science Framework is there, the infrastructure exists — and adoption is still uneven.
Iris Holm: Voluntary and uneven is the whole problem. A tenure committee reads a CV. They count publications in high-Impact-Factor journals. Preregistration is metadata. It doesn't appear in that count.
Lila Soto: So Registered Reports end up as a parallel track — the methodologically careful researchers use them, and the prestige machine keeps running on novelty beside them. And I kind of wonder if that actually makes it worse — like it gives the system a safety valve without fixing the pressure.
Iris Holm: The question is whether a tenure committee will ever count a Registered Report the same as a Nature paper. And the answer depends entirely on whether funders and hiring bodies enforce it — not whether the journals offer it.
Lila Soto: So the tools exist. The enforcement moment — the hiring meeting, the grant review — that's still running the old criteria. And until those rooms change what they reward, the 36% doesn't move.
Iris Holm: And that's the thing nobody in the reform conversation wants to say out loud. The Open Science Framework exists. Preregistration exists. Registered Reports exist. The architecture for a better system is sitting there. And a tenure committee still opens a CV and counts Nature papers. That's it. That's the whole enforcement problem in one room.
Lila Soto: Yeah. And I guess what I keep wondering — not rhetorically, like actually — what would it take for a university to genuinely hire someone whose whole career is careful replication work? Not one replication study among other things. A whole career of it. What does that candidate even look like to a hiring committee running on Journal Impact Factor?
Iris Holm: Invisible. Or — I mean, worse than invisible. They look like someone who couldn't generate their own ideas. That's the read. Even if the work is the most scientifically valuable thing produced in that department that decade.
Lila Soto: Which is kind of where I land, actually. Not with an answer — just with the weight of it. The reforms are real. Preregistration is real, Registered Reports are real, Brian Nosek built actual infrastructure at the Center for Open Science. And it's still not enough if the room where it counts never changes what it's counting.
Iris Holm: That's the uncomfortable place this leaves you. Good tools. Broken adoption. The 36% doesn't move until the hiring meeting does. I don't have a cleaner version of that.