
A study of 1,682 readers found that stories generated by ChatGPT 4.0 were rated significantly higher than human-written stories on quality and immersion—until readers learned an AI wrote them, at which point ratings dropped.
The findings show that people's judgments of creative work are shaped more by their beliefs about its origin and their attitudes toward AI than by the text itself.
Even when directly comparing human and AI stories, readers could not reliably tell them apart, though familiarity with AI systems (but not fiction expertise) helped some identify the source.
What happened
Researchers presented 1,682 participants with six short stories (about 1,000 words each)—three human-written from literary magazines, three generated by ChatGPT 4.0. Half were told each story was human-written; half were told it came from AI. ChatGPT stories scored significantly higher: mean quality score of 1.54 versus 0.97 for human stories on a scale from minus 3 to plus 3; immersion scores were 1.42 versus 1.00. Crucially, these higher ratings reversed when participants learned an AI wrote the story—their scores dropped regardless of actual authorship.
Why it matters
People's judgments of creative work are shaped less by the text itself than by who they believe wrote it and their own attitudes toward AI. Readers with positive AI attitudes gave higher ratings overall and even higher when told ChatGPT wrote the story; skeptical readers showed the opposite effect. The study suggests AI-generated text feels smoother and more emotionally accessible than literary fiction, which is often intentionally challenging—but that perceived ease does not necessarily mean better writing. For creators and publishers, this reveals a trust and perception gap: readers may enjoy AI work without believing AI capable of producing it.
What to watch
In two additional experiments with 905 participants who read both human and AI stories side-by-side and had to identify which was which, participants performed no better than chance. However, self-reported experience with AI systems correlated with the ability to correctly identify stories' origins—experience with fiction did not. The research was published in the journal Judgment and Decision Making, with all data freely available on the Open Science Framework.
Researchers Sydney Sears and Deena Skolnick Weisberg, publishing in the journal Judgment and Decision Making, conducted a series of experiments to test how people judge AI-generated versus human-written creative work. In the first experiment, 1,682 participants each read one of six short stories, approximately 1,000 words in length. Three stories came from well-known literary magazines and short story collections. The other three were generated using ChatGPT 4.0, with prompts designed to match the theme, style, and narrative perspective of the human originals.
The experiment employed a deception design: half the participants were told their assigned story was human-written, and half were told it came from ChatGPT. Critically, this information was accurate for only half of each group—meaning some readers were deliberately misled about authorship. The results showed a clear preference for AI-generated work. ChatGPT stories received a mean quality score of 1.54 compared to 0.97 for human-written stories on a scale from minus 3 to plus 3. On immersion, the gap was 1.42 for AI stories versus 1.00 for human ones. Yet when participants learned a human wrote the story, ratings in both groups rose above the levels given when the story was attributed to AI, demonstrating that the label alone powerfully shaped judgment.
Participants' attitudes toward AI also colored their responses. Those with positive attitudes toward AI consistently gave higher ratings across the board. When they were also told the story came from ChatGPT, their scores climbed even further. Participants skeptical of AI showed the opposite pattern. The researchers also examined whether people could distinguish human from AI work when given a direct comparison. In two additional experiments involving 905 total participants, each person read both a human-written and an AI-generated story and then attempted to identify which was which. Even with a side-by-side comparison, participants performed no better than chance. However, self-reported experience with AI systems correlated positively with the ability to correctly identify stories' origins, whereas self-reported experience with fiction did not improve performance.
The authors propose that AI-generated texts succeed partly because they tend to be smoother, easier to read, and more emotionally upbeat than human-written texts—traits that make material easier to process. Literary fiction, by contrast, is often intentionally difficult, designed to push readers to work for meaning. A story can be high-quality in a literary sense but still not engaging, and vice versa. The short-story format may also advantage AI; maintaining narrative coherence over 1,000 words is a different task than sustaining it across hundreds of pages. Nonetheless, the researchers' conclusion is unambiguous: AI can generate creative works that people perceive as at least equal to human work, yet people do not believe AI is capable of doing so. All data and materials from the study are freely available on the Open Science Framework. A separate study from Stony Brook University and Columbia Law School published last October adds nuance: professional readers preferred human texts with simple prompts, but when models were specifically trained on individual authors' styles, the experts preferred the AI-generated texts eight times more often for style imitation and twice as often for writing quality.
The study reveals a fundamental mismatch between what readers experience and what they believe about AI creativity. ChatGPT's stories outscored human-written texts on both quality and immersion—a finding that stands even before the labeling manipulation. Yet the moment readers learned an AI wrote the story, their ratings dropped regardless of which text they were actually judging. This suggests readers hold a conviction that AI cannot produce work as good as humans, even when their own experience contradicts that belief.
The researchers propose that AI-generated text succeeds partly because it tends to be smoother, easier to read, and emotionally upbeat—traits that make material easier to process. Literary fiction, by contrast, is often intentionally difficult, asking readers to work for meaning. Perceived ease does not equal literary quality, but it does influence how people rate what they read. The short-story format—just 1,000 words—may also favor AI, since maintaining narrative coherence is a different challenge in that length than across hundreds of pages.
A key factor is the reader's own attitude toward AI. Those with positive attitudes gave higher ratings across the board; when also told the story was AI-generated, their scores rose even further. Skeptical readers showed the opposite effect. This pattern mirrors an earlier finding in AI-generated poems: human judgment of creative work is filtered through prior belief and disposition, not just aesthetic response. Interestingly, self-reported experience with AI systems correlated with the ability to correctly identify a story's origin, while experience with fiction did not—suggesting that familiarity with how AI works matters more than literary expertise when trying to detect its hand.
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytime
Ask AI anything about this article. Q&As are published on this page for other readers too.
Amazon and Google are intensifying competitive efforts against The Trade Desk (TTD), a major digital advertisi…

OpenAI introduced Premium Seats for ChatGPT Business, priced at $125 per user per month ($100 with annual bill…

Computer scientists at University of Tübingen, Max Planck Institute, MATS Research, and Snyk discovered a meth…

Anthropic pledged to embed machine-readable watermarks in Claude-generated text and digitally signed provenanc…

Anthropic has signed the EU AI Act Code of Practice and will embed invisible watermarks in Claude-generated te…

Anthropic has agreed to pay $9.1 billion over 20 years to Riot Platforms Inc., a Bitcoin miner turned data cen…

The AI news that matters, in one minute each morning.
Sign up free