Onpode
Cover art for Study finds people prefer AI-generated stories over human-written ones and can't tell the difference

Study finds people prefer AI-generated stories over human-written ones and can't tell the difference

August 5, 2026 · 6 min

Marcus Vale & Ben Okonkwo

A Villanova University study of 1,682 participants found AI-generated stories rated 6% higher on quality and 8% more engaging than human-written ones in blind tests. Preference rose a further 3% when readers were falsely told an AI story was human-authored, suggesting the backlash against AI writing is about origin-knowledge, not reading experience.

Recent experimental research, prominently covered by New Scientist, reveals a striking gap between stated cultural attitudes toward AI-generated content and actual reader behavior.

0:006:18
Get the next episode on Artificial Intelligence

Follow it free — new episodes land in your feed.

Or make your own — any topic, in minutes

More Onpode episodes on Artificial Intelligence

About this episode

A study out of Villanova University put 1,682 readers in front of short stories — some human-written, some from ChatGPT — and asked them to rate quality and engagement. AI stories came out ahead: 6% higher on quality, 8% more engaging. A separate replication in the Journal of Judgment and Decision Making found the same pattern across adults from 18 to 81. The episode takes that result seriously and follows it somewhere uncomfortable. If readers preferred AI stories blind, but preference actually increases when an AI story is labeled as human-written, then the backlash isn't about the reading experience at all — it's an origin-knowledge effect. The problem is that readers can identify AI authorship at roughly coin-flip accuracy (40–52%), which means the cultural penalty for AI and the authenticity bonus for human writing are both downstream of a judgment most people demonstrably cannot make. The episode also catches something the BBC and Time coverage missed: stylometric analysis using Burrows' Delta can statistically fingerprint GPT-3.5, GPT-4, and Llama outputs even when human readers fail entirely. 'Undetectable' is doing a lot of work in this conversation, and it hasn't been earned. The episode ends on a question nobody in publishing seems eager to answer: if you already have the data showing readers rate the same content higher without disclosure, what exactly is the incentive to keep the label on?

Frequently asked

Do people prefer AI-written stories over human-written stories?

Yes. A Villanova University study led by Haoran Chu, published in the Journal of Communication, found 1,682 participants rated ChatGPT-generated stories 6% higher on quality and 8% more engaging than human-written stories in blind tests — without being able to tell the difference.

Can people tell the difference between AI-written and human-written stories?

No, not reliably. A study published in the journal Judgment and Decision Making found adults aged 18 to 81 detected AI authorship at roughly 40–52% accuracy — statistically no better than a coin flip. However, stylometric algorithms like Burrows' Delta can distinguish GPT-3.5, GPT-4, and Llama 70b outputs even when human readers cannot.

Does knowing a story was written by AI change how people rate it?

Yes, but in a surprising direction. The Villanova study found readers' preference for AI stories increased by an additional 3% when they were told an AI-written story was actually human-authored. The penalty for AI origin only activates when readers know the true source — and they usually cannot determine that reliably.

Can AI-generated text be detected by algorithms even if humans can't spot it?

Yes. Stylometric analysis using Burrows' Delta finds statistically distinguishable patterns in GPT-3.5, GPT-4, and Llama 70b outputs, even when human readers perform at chance level. Counterintuitively, GPT-4 was more detectable by the algorithm than GPT-3.5, though whether newer models will close that gap remains unsettled.

Why does the AI writing backlash exist if readers actually prefer AI stories?

The backlash appears to be an origin-knowledge effect, not an aesthetic one. Research shows readers rate identical text higher when they believe it is human-authored, and the anti-AI penalty only fires when authorship is disclosed. Because readers can't reliably detect AI writing, the penalty is applied inconsistently or not at all.

Grounded in 8 sources
Human-AI Interaction Traces as Blackout Poetry: Reframing AI-Supported Writing as Found-Text Creativity · arxiv.org
Bot or not: Can people tell the difference between stories written by a human or by an AI system? | Judgment and Decision Making | Cambridge Core · cambridge.org
Detection of AI versus human-authored narratives · nature.com
Stylometric comparisons of human versus AI-generated creative writing | Humanities and Social Sciences Communications · nature.com
Ability of AI detection tools and humans to accurately identify different forms of AI-generated written content · pmc.ncbi.nlm.nih.gov
Of love & lasers: Perceptions of narratives by AI versus human authors · sciencedirect.com
AI-generated tales preferred to human-written ones, study finds - BBC News · bbc.co.uk
AI-generated tales preferred to human-written ones, study finds · bbc.com
Read transcript

Ben Okonkwo: Marcus, hey — I had a weird moment this week. Someone in my reading group passed around a short story, asked everyone to rate it. Genuinely moving, people loved it. Turned out it was ChatGPT. The group's response was immediate outrage — but their ratings were already in.

Marcus Vale: That's the whole study, basically.

Ben Okonkwo: It is — which is what made it so strange to watch live. Because Haoran Chu ran exactly that experiment at Villanova University, 1,682 participants, published in the Journal of Communication. ChatGPT stories rated six percent higher on quality, eight percent more engaging than human-written ones. Blind.

Marcus Vale: And then the Judgment and Decision Making journal replicates it — adults 18 to 81, same direction. So the backlash. Where is it actually coming from if it's not coming from the reading experience?

Ben Okonkwo: Interesting — because I think the three-percent label effect is the clue there. Preference actually increased when readers were told an AI story was human-authored. So the aversion isn't aesthetic. It's... origin-knowledge triggering something else entirely.

Marcus Vale: Identity signal. Not quality signal. New Scientist flagged this as directly contradicting the cultural backlash narrative — and I think they're right, the data just doesn't leave much room.

Ben Okonkwo: Which means authenticity isn't a property of the text — it's a label you stick on it. Think about the wine thing: someone hands you two glasses, no labels, you pick one. You love it. Then they tell you the one you loved was the eight-dollar bottle. Your experience was real. Your belief about origin was completely separate from that experience. That's the whole mechanism in one sentence. The Villanova study found that additional three percent preference lift not from the writing changing — from the label changing.

Marcus Vale: Hold on — so the anti-AI penalty only fires when you actually know it's AI.

Ben Okonkwo: Right — the problem is if detection is running at forty to fifty-two percent, which is basically coin-flip territory from the Judgment and Decision Making study, readers can't reliably apply that penalty. They don't know when to dock points. So the whole preference architecture — the backlash, the origin bonus, all of it — is sitting on top of a label they have no reliable way to verify.

Marcus Vale: That's the trap door. The authenticity bonus and the AI penalty are both downstream of a judgment readers demonstrably cannot make. So imagine a book club member — she's just finished rating a story 'deeply human' out loud in the discussion. Confident. Then someone shows her it was ChatGPT. The rating was real. The label was wrong. The experience was identical. And she had no tools to catch it.

Ben Okonkwo: Which is exactly what happened in your reading group anecdote from the top — the outrage came after the ratings were already locked in.

Marcus Vale: So 'undetectable' becomes a market position, not a technical claim. Which is why, frankly, the disclosure question gets really interesting — the EU rules kicking in August 2026 are trying to restore a signal readers can't generate on their own.

Ben Okonkwo: Now — and this is worth flagging because we'll get there — stylometry, Burrows' Delta applied to GPT-3.5 and GPT-4 outputs, finds statistically distinguishable patterns even when human readers completely miss them. So 'undetectable' is not actually a settled claim. That part is more complicated than it looks.

Marcus Vale: And that's the crack I want to pull on — because 'undetectable' is doing a lot of work in this conversation and it's not actually earned. Burrows' Delta stylometric analysis finds statistically distinguishable patterns across GPT-3.5, GPT-4, and Llama 70b, even when human readers are at coin-flip accuracy. So the question isn't whether detection is possible. It's whether we're asking the right instrument.

Ben Okonkwo: Right — and that's a meaningful distinction. Human readers at chance level is not the same claim as 'undetectable.' Those are two completely different statements.

Marcus Vale: GPT-4 is actually more distinguishable than GPT-3.5 under Burrows' Delta. Which is — wait, that's the counterintuitive part. The more capable model is easier for the algorithm to fingerprint, not harder.

Ben Okonkwo: Hm — okay, I want to flag something there, because I'm not sure the research settles what that means going forward. GPT-3.5 showed occasional stylistic overlap with human writing. GPT-4 was more distinguishable. That could mean detection gets easier as models evolve, or it could mean the next generation closes that gap in a completely different direction. The arms race framing is real, but I don't think the data tells us which way it resolves.

Marcus Vale: Fair. But BBC and Time both covered the Villanova findings without touching the stylometry counterpoint once. That's the story getting flattened in real time.

Ben Okonkwo: Now — the Nature cross-cultural study is actually where I think the mechanism gets interesting. Aesthetics-trained participants in Chinese contexts, working collaboratively, could identify AI text by focusing on cultural depth. Not surface style. What the algorithm catches statistically, trained human groups can catch culturally, but only when they're looking at the right thing.

Marcus Vale: So detection isn't a fixed cognitive limit — it's a skill gap. Which means the real bet is whether that gap closes through tools, training, or both before the content ecosystem fully prices in 'undetectable' as a given.

Ben Okonkwo: The thing that I keep sitting with — the disclosure question. If preference for AI stories rises when readers are told they're human-authored, then the label isn't restoring authenticity. It's creating a penalty that wasn't in the reading experience itself. The experience was already better without it.

Marcus Vale: Which means — if publishers already have the Villanova data, if they know readers rate the same content higher without disclosure, what's the actual incentive to keep the label on?

Ben Okonkwo: I don't have a good answer to that.

Marcus Vale: Neither do I. Good talk.