AIToday
Large Language ModelsTHE DECODERPublished: Aug 9, 2026, 01:01 JST6 min read

AI stories rated higher than human ones—until readers learn the truth

AI stories rated higher than human ones—until readers learn the truth

Key takeaway

  • A study of 1,682 readers found that stories generated by ChatGPT 4.0 were rated significantly higher than human-written stories on quality and immersion—until readers learned an AI wrote them, at which point ratings dropped.

  • The findings show that people's judgments of creative work are shaped more by their beliefs about its origin and their attitudes toward AI than by the text itself.

  • Even when directly comparing human and AI stories, readers could not reliably tell them apart, though familiarity with AI systems (but not fiction expertise) helped some identify the source.

3 Key Points

  1. What happened

    Researchers presented 1,682 participants with six short stories (about 1,000 words each)—three human-written from literary magazines, three generated by ChatGPT 4.0. Half were told each story was human-written; half were told it came from AI. ChatGPT stories scored significantly higher: mean quality score of 1.54 versus 0.97 for human stories on a scale from minus 3 to plus 3; immersion scores were 1.42 versus 1.00. Crucially, these higher ratings reversed when participants learned an AI wrote the story—their scores dropped regardless of actual authorship.

  2. Why it matters

    People's judgments of creative work are shaped less by the text itself than by who they believe wrote it and their own attitudes toward AI. Readers with positive AI attitudes gave higher ratings overall and even higher when told ChatGPT wrote the story; skeptical readers showed the opposite effect. The study suggests AI-generated text feels smoother and more emotionally accessible than literary fiction, which is often intentionally challenging—but that perceived ease does not necessarily mean better writing. For creators and publishers, this reveals a trust and perception gap: readers may enjoy AI work without believing AI capable of producing it.

  3. What to watch

    In two additional experiments with 905 participants who read both human and AI stories side-by-side and had to identify which was which, participants performed no better than chance. However, self-reported experience with AI systems correlated with the ability to correctly identify stories' origins—experience with fiction did not. The research was published in the journal Judgment and Decision Making, with all data freely available on the Open Science Framework.

In Depth

Read the full story

Researchers Sydney Sears and Deena Skolnick Weisberg, publishing in the journal Judgment and Decision Making, conducted a series of experiments to test how people judge AI-generated versus human-written creative work. In the first experiment, 1,682 participants each read one of six short stories, approximately 1,000 words in length. Three stories came from well-known literary magazines and short story collections. The other three were generated using ChatGPT 4.0, with prompts designed to match the theme, style, and narrative perspective of the human originals.

The experiment employed a deception design: half the participants were told their assigned story was human-written, and half were told it came from ChatGPT. Critically, this information was accurate for only half of each group—meaning some readers were deliberately misled about authorship. The results showed a clear preference for AI-generated work. ChatGPT stories received a mean quality score of 1.54 compared to 0.97 for human-written stories on a scale from minus 3 to plus 3. On immersion, the gap was 1.42 for AI stories versus 1.00 for human ones. Yet when participants learned a human wrote the story, ratings in both groups rose above the levels given when the story was attributed to AI, demonstrating that the label alone powerfully shaped judgment.

Participants' attitudes toward AI also colored their responses. Those with positive attitudes toward AI consistently gave higher ratings across the board. When they were also told the story came from ChatGPT, their scores climbed even further. Participants skeptical of AI showed the opposite pattern. The researchers also examined whether people could distinguish human from AI work when given a direct comparison. In two additional experiments involving 905 total participants, each person read both a human-written and an AI-generated story and then attempted to identify which was which. Even with a side-by-side comparison, participants performed no better than chance. However, self-reported experience with AI systems correlated positively with the ability to correctly identify stories' origins, whereas self-reported experience with fiction did not improve performance.

The authors propose that AI-generated texts succeed partly because they tend to be smoother, easier to read, and more emotionally upbeat than human-written texts—traits that make material easier to process. Literary fiction, by contrast, is often intentionally difficult, designed to push readers to work for meaning. A story can be high-quality in a literary sense but still not engaging, and vice versa. The short-story format may also advantage AI; maintaining narrative coherence over 1,000 words is a different task than sustaining it across hundreds of pages. Nonetheless, the researchers' conclusion is unambiguous: AI can generate creative works that people perceive as at least equal to human work, yet people do not believe AI is capable of doing so. All data and materials from the study are freely available on the Open Science Framework. A separate study from Stony Brook University and Columbia Law School published last October adds nuance: professional readers preferred human texts with simple prompts, but when models were specifically trained on individual authors' styles, the experts preferred the AI-generated texts eight times more often for style imitation and twice as often for writing quality.

Context & Analysis

The study reveals a fundamental mismatch between what readers experience and what they believe about AI creativity. ChatGPT's stories outscored human-written texts on both quality and immersion—a finding that stands even before the labeling manipulation. Yet the moment readers learned an AI wrote the story, their ratings dropped regardless of which text they were actually judging. This suggests readers hold a conviction that AI cannot produce work as good as humans, even when their own experience contradicts that belief.

The researchers propose that AI-generated text succeeds partly because it tends to be smoother, easier to read, and emotionally upbeat—traits that make material easier to process. Literary fiction, by contrast, is often intentionally difficult, asking readers to work for meaning. Perceived ease does not equal literary quality, but it does influence how people rate what they read. The short-story format—just 1,000 words—may also favor AI, since maintaining narrative coherence is a different challenge in that length than across hundreds of pages.

A key factor is the reader's own attitude toward AI. Those with positive attitudes gave higher ratings across the board; when also told the story was AI-generated, their scores rose even further. Skeptical readers showed the opposite effect. This pattern mirrors an earlier finding in AI-generated poems: human judgment of creative work is filtered through prior belief and disposition, not just aesthetic response. Interestingly, self-reported experience with AI systems correlated with the ability to correctly identify a story's origin, while experience with fiction did not—suggesting that familiarity with how AI works matters more than literary expertise when trying to detect its hand.

FAQ

How many people did the study include?
The first experiment included 1,682 participants. Two additional experiments with a direct comparison task involved 905 total participants.
Did people actually prefer the AI stories, or was it just about the label?
ChatGPT stories scored significantly higher on both quality (1.54 versus 0.97) and immersion (1.42 versus 1.00) even before the labeling manipulation. However, when participants were told a human wrote the story, ratings in both groups rose higher than when the story was attributed to AI—showing that the label strongly shaped perception.
Could people tell human and AI stories apart?
No. In the experiments where participants read both a human-written and AI-generated story and had to identify which was which, they performed no better than chance. Self-reported experience with AI systems did correlate with the ability to correctly identify the stories' origins, but experience with fiction did not help.

Get the latest Large Language Models news every morning

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytime

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleSnowflake cuts contract review time 70% with AI agent

The AI news that matters, in one minute each morning.

Sign up free