AIToday
Large Language ModelsRoboticsHacker NewsPublished: Jul 16, 2026, 01:00 JST3 min read

AI detectors flag human writing from before ChatGPT as mostly machine

AI detectors flag human writing from before ChatGPT as mostly machine

Key takeaway

  • A writer tested eleven AI-detection tools on her own provably human writing spanning 2004–2026, including texts predating ChatGPT by years. The detectors disagreed wildly—flagging the same 2019 text as anywhere from 5% to 100% human—and systematically misclassified her pre-LLM archive as mostly machine-written.

  • Three tools failed the control group entirely, marking all eight of her compositions as majority AI.

  • The detectors appear to mistake her particular writing style (structured, high information density) for the hallmarks of AI rather than measuring actual authorship, undermining their use by educators and publishers to reject submissions and accuse writers of fraud.

3 Key Points

  1. What happened

    A writer tested eleven AI-detection tools on eight pieces of her own writing, including four texts written before ChatGPT existed (2019–2020). Results ranged wildly: one tool flagged her 2019 setup guide as only 5% human, while six others scored it 100% human. Her grad school essay from 2020 scored 15% human; a social media piece scored 12%.

  2. Why it matters

    These detectors are being used by teachers, editors, and publishers to accuse real people of fraud—Clarkesworld and Asimov's Science Fiction both reject AI-assisted submissions, with risk of permanent bans. Yet the tools disagree so sharply on the same text (a 95-point spread on her setup guide) that they are unreliable as instruments. Three detectors failed the control group outright, marking all eight of her human pieces as majority AI, including pre-LLM writing.

  3. What to watch

    The writer's most AI-assisted work (a serialized fiction project where Claude helps with continuity, character sheets, and research) scored 99–100% human on seven of eleven detectors—higher than her pre-LLM archive. The detectors appear to measure predictability and writing style (high lexical density, structured prose) rather than authorship, flagging her particular voice as suspicious regardless of era.

Ask the AI about this article →

Context & Analysis

The core problem is that AI-detection tools are not measuring authorship—they are measuring predictability and stylistic consistency. The writer's controlled experiment exposes a fundamental flaw: three detectors (Ace, Humalingo, and Getsolved) flagged all eight of her human-written pieces, including texts from 2004–2020 that predate ChatGPT by years, as majority or near-entirely AI-generated. This is not a marginal disagreement; it is a categorical failure. Two thermometers reading the same winter day as 5 and 100 degrees do not both measure temperature—one is broken and useless.

The detectors appear to flag her writing because of its structural features: high lexical density, strong syntactic consistency, and information precision. These are hallmarks of her deliberate, systems-thinking voice—and they happen to mirror the output of language models running at low temperature, where less probable tokens are suppressed. But low-temperature sampling is not authorship; it is one technical mechanism among many. Her most AI-assisted piece, where Claude handled filing, continuity, and research, scored as more human than her pre-LLM archive on most detectors, because it contained unpredictable narrative choices (gravity as a mob boss, the strong nuclear force as a seductress) that fall outside the detectors' training distributions.

The stakes are concrete and immediate: Clarkesworld and Asimov's Science Fiction both reject AI-assisted work, with permanent bans for attempted submissions. Teachers are using these tools to accuse students of fraud. Editors and publishers are rejecting submissions. The writer's evidence suggests these institutions are deploying instruments that cannot agree with each other and that systematically misclassify human writing created before the technology existed.

FAQ

How many AI detectors were tested and what were the most extreme disagreements?
Eleven detectors were tested: Originality, Ace, Humalingo, ZeroGPT, Grammarly, GPTZero, Getsolved, Quillbot, Pangram, CopyLeaks, and WinstonAI. Her 2019 setup guide scored 5% human according to Getsolved but 100% human according to six other tools—a 95-point spread on the same words.
Which publications are banning AI-assisted submissions?
Clarkesworld (July 2026) and Asimov's Science Fiction (July 2026) both reject submissions written, developed, or assisted by AI tools, with Asimov's warning that attempted submissions may result in permanent bans from future submissions.
What was the writer's most AI-assisted work and how did detectors score it?
A serialized fiction project about physics and vampires, where Claude helps maintain continuity, character sheets, and research. It scored 99–100% human on seven of eleven detectors, 91% on an eighth, and earned the highest average humanity score of all eight pieces tested—higher than her pre-LLM archive.

Get the latest Large Language Models news every morning

For example, today's edition would include:

  • DataAgent launches with $10M to auto-fix Kubernetes faultsSiliconANGLE AI · 49m ago
  • SK Hynix custom HBM boosts inference up to 5.15xDIGITIMES Asia · 49m ago
  • Nvidia Earnings: Boring by Design, Avoiding a Consolidated WorldStratechery (Ben Thompson) · 49m ago

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleMeta infrastructure chief: 20 months to rebuild for AI agents