Eleanor Crane: You know what I noticed reading through this — I never once asked myself whether the scientists were honest. I just kept asking what they were being paid to do.
Ben Okonkwo: That's — yeah, that's the entry point, actually. Because the data doesn't flatter anyone, but it also doesn't accuse anyone. A hundred published psychology studies. The Open Science Collaboration replicates all of them. Thirty-six hold up.
Eleanor Crane: Sixty-four published findings that simply don't survive a second look.
Ben Okonkwo: And I want to flag the assumption the whole crisis narrative rests on — because some researchers, Stroebe and Strack among them, pushed back and said failed replications might just mean we missed a moderator, some contextual variable the replication didn't match. Which is not nothing.
Eleanor Crane: And is that a live possibility, or does that number — sixty-four out of a hundred — make the moderator explanation feel like a rescue?
Ben Okonkwo: Hm. I think it's both, and the ambiguity matters. But here's what the structural story adds that the moderator story can't: journals don't publish null results. Hiring committees don't reward replications. So even if some failures reflect missing moderators, the system has no mechanism for distinguishing those from genuine false positives — and no incentive to try.
Eleanor Crane: So the real question isn't whether individual papers are wrong — it's how science built a pipeline that can't tell the difference.
Ben Okonkwo: That's what today is actually about, I think — how a system gets built to reward exactly the behavior that breaks it.
Eleanor Crane: And that pipeline — I mean, the way it actually manufactures a false positive isn't dramatic at all. It's almost boring. Imagine someone buys twenty lottery tickets, wins once, and writes a paper titled 'I Have a System.' That's p-hacking in a sentence. You keep adjusting the analysis, shifting which variables you include, peeking at the results, until the number you need finally appears — and then you write up that version as if it were the only version you ever ran.
Ben Okonkwo: And you never mention the nineteen losing tickets.
Eleanor Crane: You never mention the nineteen losing tickets.
Ben Okonkwo: Now, there's a second move that makes this worse — actually, I'd say it's what turns a lucky result into a convincing-looking prediction. You run the analysis, you see what came out, and then you write the hypothesis as if you'd had it before you looked. HARKing — hypothesizing after results are known. It's like watching a dart hit a random spot on the wall and then walking up and drawing the bullseye around it. To any reader, it looks like you called your shot.
Eleanor Crane: Which is — wait, that's not even detectable from the outside, is it? The paper just reads as a clean prediction.
Ben Okonkwo: That's exactly the problem. And then — okay, compound it further. Run enough separate tests on the same dataset and a false positive becomes nearly guaranteed just by chance. Not by intent, not by bad faith. Mathematically inevitable. That's the multiple comparisons problem. The more questions you ask of the data, the more the data will eventually hand you something that looks like an answer.
Eleanor Crane: So the stated significance level — the number that's supposed to tell you how confident to be — is already optimistic before anyone does anything deliberately wrong.
Ben Okonkwo: Right — and when you stack p-hacking on top of HARKing on top of multiple comparisons, the false-positive rate inflation isn't a rounding error. It's the whole ballgame. Published 'significant' findings are carrying far less signal than the label says. And that's before a single result enters the file drawer and disappears because it didn't land.
Eleanor Crane: And the file drawer is where this compounds. Because it's not just that the false positives get published — it's that the null results, the ones that would balance the picture, they just... vanish. Someone runs the study, gets nothing, and puts it in a folder. The journal wouldn't take it anyway.
Ben Okonkwo: Right — and that asymmetry is what canonizes the error. Ego-depletion is the case I keep returning to. Roy Baumeister's paradigm. Earlier meta-analyses looked at the published literature and saw confirmation. Apparent consensus.
Eleanor Crane: A graduate student in 2010 reads those meta-analyses and thinks — the field has settled this. Builds a dissertation on it.
Ben Okonkwo: And the consensus was real, in the sense that it was in the literature. But the literature was — I mean, it was a biased sample. The publication-bias-corrected meta-analyses, starting around 2014, found no statistically reliable evidence for ego-depletion at all. Once you account for the file drawer, the effect disappears.
Eleanor Crane: The literature itself was lying. Not any single paper — the whole accumulated structure.
Ben Okonkwo: Now take that pattern into cancer biology, and the stakes shift completely. Begley and Ellis, 2012 — they tried to confirm fifty-three pre-clinical cancer studies. Landmark findings, top journals. Eleven percent replicated. Eleven.
Eleanor Crane: Wait — eleven percent of fifty-three. That's... six studies.
Ben Okonkwo: Roughly. And then the Reproducibility Project: Cancer Biology went further — fifty-three top cancer papers, 2010 to 2012, one hundred and ninety-three experiments total. Fifty experiments replicated. And not one of those original papers had fully described its experimental protocols. Seventy percent of replications required contacting authors directly just to get the reagents.
Eleanor Crane: Clinicians were reading those papers. That's the part I can't quite — and what comes next, whether the structural incentives can actually be reformed, that's even harder to sit with.
Ben Okonkwo: And that's the part that doesn't resolve cleanly — because the structural incentive isn't just indifferent to replication, it actively punishes it. Replication studies face publication bias against null results. There's no meaningful career reward for running one. The corrective mechanism has no professional engine.
Eleanor Crane: Which means the system can diagnose its own errors and still — what, shrug them off? That's not a bug someone forgot to fix.
Ben Okonkwo: Now, here's the sharpest version of this I've seen — there's economics PhD job market research showing that marginally significant results were associated with better academic placements than robust, solid ones. Better jobs. The researchers who were closest to the p-hacking line got rewarded for it.
Eleanor Crane: Wait — better placements for marginally significant results than for clean ones?
Ben Okonkwo: Better placements. Hiring committees, implicitly, valued the appearance of a surprising finding over a reliable one. That's the system working as designed — it's just designed wrong.
Eleanor Crane: So the question I keep turning over is — can you reform the mechanism without touching what the mechanism rewards? And people have tried. Registered Reports, which — I mean, PLOS Biology adopted this, Royal Society Open Science — the journal peer-reviews the design before you collect any data, provisionally accepts it, so the result can't be buried regardless of what comes out.
Ben Okonkwo: And the Center for Open Science built the whole scaffolding around this — Open Science Framework, the TOP Guidelines, pre-registration registries where you time-stamp your hypotheses before you run anything. In principle, it closes the HARKing loop entirely.
Eleanor Crane: In principle. But Campbell's Law is — actually, that's the critique that won't leave me alone. When a measure becomes a target, it loses its validity as a measure. Pre-registration becomes a badge. And badges get performed.
Ben Okonkwo: And Stroebe and Strack are still waiting in the wings — their moderator defense becomes the institutional antibody. Failed replication? Unmeasured contextual variable. Nothing to see. It insulates the original finding without engaging the evidence. Two layers of protection and the incentive structure never shifts.
Eleanor Crane: So if hiring committees still count Nature papers and journals still chase novelty — does pre-registration change the behavior, or just the paperwork?
Ben Okonkwo: Just the paperwork, I think. And that's — I mean, I can't shake the 36-out-of-100 number, not because it's a scandal about dishonest researchers, but because if you model it: journals reward novelty, hiring committees reward novelty, replication carries no professional engine — then 36 isn't shocking. It's almost the predicted output. A researcher can build a career on a novel false positive. Publishing a true negative costs years. Those aren't equivalent choices.
Eleanor Crane: And yet someone ran those replications. The Open Science Collaboration got a hundred psychologists to do work that had no career upside. That's the thing I can't quite place inside your model — not as a happy ending, but as a fact that doesn't fit the incentive story cleanly.
Ben Okonkwo: No, you're right. And I genuinely don't know what to do with that — whether it means the values and the incentives can diverge long enough to matter, or whether it's just a one-time coordination that won't hold without something structural underneath it.
Eleanor Crane: That's where I keep getting stuck too. What would have to change in what academia actually values — not what it requires on a checklist, but what it rewards when no one's watching — for accuracy to outrank surprise. I don't have an answer to that.
Ben Okonkwo: Neither do I. And I think that's probably the honest place this lands. Thanks for thinking through it with me.