
What happened
Researchers presented 1,682 participants with six short stories (about 1,000 words each)—three human-written from literary magazines, three generated by ChatGPT 4.0. Half were told each story was human-written; half were told it came from AI. ChatGPT stories scored significantly higher: mean quality score of 1.54 versus 0.97 for human stories on a scale from minus 3 to plus 3; immersion scores were 1.42 versus 1.00. Crucially, these higher ratings reversed when participants learned an AI wrote the story—their scores dropped regardless of actual authorship.
Why it matters
People's judgments of creative work are shaped less by the text itself than by who they believe wrote it and their own attitudes toward AI. Readers with positive AI attitudes gave higher ratings overall and even higher when told ChatGPT wrote the story; skeptical readers showed the opposite effect. The study suggests AI-generated text feels smoother and more emotionally accessible than literary fiction, which is often intentionally challenging—but that perceived ease does not necessarily mean better writing. For creators and publishers, this reveals a trust and perception gap: readers may enjoy AI work without believing AI capable of producing it.
What to watch
In two additional experiments with 905 participants who read both human and AI stories side-by-side and had to identify which was which, participants performed no better than chance. However, self-reported experience with AI systems correlated with the ability to correctly identify stories' origins—experience with fiction did not. The research was published in the journal Judgment and Decision Making, with all data freely available on the Open Science Framework.
Summaries like this, in your inbox every morning.
The study reveals a fundamental mismatch between what readers experience and what they believe about AI creativity. ChatGPT's stories outscored human-written texts on both quality and immersion—a finding that stands even before the labeling manipulation. Yet the moment readers learned an AI wrote the story, their ratings dropped regardless of which text they were actually judging. This suggests readers hold a conviction that AI cannot produce work as good as humans, even when their own experience contradicts that belief.
The researchers propose that AI-generated text succeeds partly because it tends to be smoother, easier to read, and emotionally upbeat—traits that make material easier to process. Literary fiction, by contrast, is often intentionally difficult, asking readers to work for meaning. Perceived ease does not equal literary quality, but it does influence how people rate what they read. The short-story format—just 1,000 words—may also favor AI, since maintaining narrative coherence is a different challenge in that length than across hundreds of pages.
A key factor is the reader's own attitude toward AI. Those with positive attitudes gave higher ratings across the board; when also told the story was AI-generated, their scores rose even further. Skeptical readers showed the opposite effect. This pattern mirrors an earlier finding in AI-generated poems: human judgment of creative work is filtered through prior belief and disposition, not just aesthetic response. Interestingly, self-reported experience with AI systems correlated with the ability to correctly identify a story's origin, while experience with fiction did not—suggesting that familiarity with how AI works matters more than literary expertise when trying to detect its hand.
For example, today's edition would include:
AI-summarized, only the topics you pick: one digest a day via Email, LINE, or Slack.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.
Autoheal AI Inc. raised $7.9 million in seed funding led by Innovation Endeavors, with Emergent Ventures, U&I…
Paul Cheek's AI-Driven Enterprise Institute study found just over 30% of S&P 500 executives are AI-literate, a…

From 7/30 to 9/17, /code-review ran 23 times with at most 1 subagent; from 9/23 it launched 10 at once, hittin…

Mizushima (technology evangelist at Nextbeat) gave Claude Fable 5.1 a five-step goal chain; it first shipped a…

A Zenn article narrowed agent cost design to three topics: cache depends on prefix stability, routing should b…

Working alone with 10 parallel Claude Code sessions, he logged 2,848 commits, 1,212 pull requests and 1,138 me…
