AIToday
Large Language ModelsTHE DECODERPublished: Aug 9, 2026, 01:01 JST

AI stories rated higher than human ones—until readers learn the truth

AI stories rated higher than human ones—until readers learn the truth

3 Key Points

  1. What happened

    Researchers presented 1,682 participants with six short stories (about 1,000 words each)—three human-written from literary magazines, three generated by ChatGPT 4.0. Half were told each story was human-written; half were told it came from AI. ChatGPT stories scored significantly higher: mean quality score of 1.54 versus 0.97 for human stories on a scale from minus 3 to plus 3; immersion scores were 1.42 versus 1.00. Crucially, these higher ratings reversed when participants learned an AI wrote the story—their scores dropped regardless of actual authorship.

  2. Why it matters

    People's judgments of creative work are shaped less by the text itself than by who they believe wrote it and their own attitudes toward AI. Readers with positive AI attitudes gave higher ratings overall and even higher when told ChatGPT wrote the story; skeptical readers showed the opposite effect. The study suggests AI-generated text feels smoother and more emotionally accessible than literary fiction, which is often intentionally challenging—but that perceived ease does not necessarily mean better writing. For creators and publishers, this reveals a trust and perception gap: readers may enjoy AI work without believing AI capable of producing it.

  3. What to watch

    In two additional experiments with 905 participants who read both human and AI stories side-by-side and had to identify which was which, participants performed no better than chance. However, self-reported experience with AI systems correlated with the ability to correctly identify stories' origins—experience with fiction did not. The research was published in the journal Judgment and Decision Making, with all data freely available on the Open Science Framework.

Not sure about something? Ask the AI

Questions and answers are published on this page.

Summaries like this, in your inbox every morning.

Context & Analysis

The study reveals a fundamental mismatch between what readers experience and what they believe about AI creativity. ChatGPT's stories outscored human-written texts on both quality and immersion—a finding that stands even before the labeling manipulation. Yet the moment readers learned an AI wrote the story, their ratings dropped regardless of which text they were actually judging. This suggests readers hold a conviction that AI cannot produce work as good as humans, even when their own experience contradicts that belief.

The researchers propose that AI-generated text succeeds partly because it tends to be smoother, easier to read, and emotionally upbeat—traits that make material easier to process. Literary fiction, by contrast, is often intentionally difficult, asking readers to work for meaning. Perceived ease does not equal literary quality, but it does influence how people rate what they read. The short-story format—just 1,000 words—may also favor AI, since maintaining narrative coherence is a different challenge in that length than across hundreds of pages.

A key factor is the reader's own attitude toward AI. Those with positive attitudes gave higher ratings across the board; when also told the story was AI-generated, their scores rose even further. Skeptical readers showed the opposite effect. This pattern mirrors an earlier finding in AI-generated poems: human judgment of creative work is filtered through prior belief and disposition, not just aesthetic response. Interestingly, self-reported experience with AI systems correlated with the ability to correctly identify a story's origin, while experience with fiction did not—suggesting that familiarity with how AI works matters more than literary expertise when trying to detect its hand.

FAQ
How many people did the study include?
The first experiment included 1,682 participants. Two additional experiments with a direct comparison task involved 905 total participants.
Did people actually prefer the AI stories, or was it just about the label?
ChatGPT stories scored significantly higher on both quality (1.54 versus 0.97) and immersion (1.42 versus 1.00) even before the labeling manipulation. However, when participants were told a human wrote the story, ratings in both groups rose higher than when the story was attributed to AI—showing that the label strongly shaped perception.
Could people tell human and AI stories apart?
No. In the experiments where participants read both a human-written and AI-generated story and had to identify which was which, they performed no better than chance. Self-reported experience with AI systems did correlate with the ability to correctly identify the stories' origins, but experience with fiction did not help.

Get the latest Large Language Models news every morning

For example, today's edition would include:

  • Autoheal raises $7.9 million for self-fixing AI agentsSiliconANGLE AI · 53m ago
  • Paul Cheek: 30% of S&P 500 execs AI-literate, 78% gapFortune AI · 53m ago
  • Claude Fable 5.1 builds matrix-free Transformer site, then a 10M Japanese SLMZenn AI/ML · 53m ago

AI-summarized, only the topics you pick: one digest a day via Email, LINE, or Slack.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.

Questions and answers are published on this page.

Related Articles

Next articleSnowflake cuts contract review time 70% with AI agent