AIToday
The Verge AIPublished: Aug 9, 2026, 22:00 JST6 min read

AI detectors fuel accusations, but keep getting it wrong

AI detectors fuel accusations, but keep getting it wrong

Key takeaway

  • Schools and online platforms are adopting AI detection tools to identify student work and articles written with generative AI, but these tools are producing significant false positives—especially against non-native English speakers—leading to wrongful accusations that have damaged writers' careers and students' academic records.

  • Despite vendors claiming low error rates, multiple universities and even OpenAI have abandoned these tools, citing poor accuracy, and educators are now pivoting toward classroom-based assessments and transparent disclosure policies instead.

3 Key Points

  1. What happened

    Schools and platforms are rapidly adopting AI detection tools to catch students and writers using generative AI, but these detectors—including GPTZero, Pangram, and Turnitin's built-in feature—are flagging human-written work as AI-generated, sometimes with serious consequences. A 2023 Stanford study found that AI detectors falsely flagged essays by non-native English speakers more often than those by native speakers. High-profile cases include publisher Minotaur dropping a $2 million book deal over AI-use concerns and a Yale student suing after failing a final exam based on an AI detection scan.

  2. Why it matters

    The tools work by analyzing writing patterns like word choice, rhythm, and tone rather than direct comparison to known sources—a method that's far less reliable than traditional plagiarism detection. Multiple educators, universities (including Yale, MIT, and Johns Hopkins), and tool makers themselves have warned the detectors are unreliable; OpenAI shut down its own AI detector in 2023 due to low accuracy. This uncertainty is creating a climate where accusations spread faster than facts, harming writers' and students' reputations and livelihoods even when they're innocent.

  3. What to watch

    Some institutions are backing away from detection tools altogether, instead asking professors to hold in-class assessments and allow students to disclose AI use without penalty. Meanwhile, platforms like Substack and LinkedIn are embedding detection buttons into their apps, likely amplifying false accusations. The Authors Guild is offering "Human Authored" certifications to help writers defend themselves.

In Depth

Read the full story

The hunt for AI-generated content has escalated from the classroom into a broader culture of accusation. For decades, educators and editors relied on anti-plagiarism tools like Turnitin to catch copied work by comparing submissions against databases of web content and published articles. These tools offer a percentage match, though they have their own flaws—false positives and uncertainty about intent. When ChatGPT, Google Gemini, and Microsoft Copilot arrived, teachers adopted a new generation of detectors just as quickly. A survey from the Center for Democracy and Technology found that 43 percent of sixth to 12th grade teachers in the US regularly used AI detectors between 2024 and 2025.

Unlike traditional plagiarism detection, AI detectors like GPTZero, Pangram, and Turnitin's built-in feature do not match text against a known database. Instead, they rely on proprietary AI models to guess whether writing is human-made by analyzing patterns in wording, rhythm, structure, and tone. Vendors claim low false positive rates: Turnitin says it falsely flags less than 1 percent of human-written content as AI; Pangram claims 1 in 10,000; GPTZero says similarly. Yet these claims have not held up in practice. A 2023 Stanford study found that AI detectors falsely flagged essays written by non-native English speakers as AI more often than native speakers—a critical failure that contradicts vendor assurances. The detectors may also be biased against neurodivergent writers. Tools pick up on patterns like repetitive terms, overly formal or informal tone, consistent sentence structure, and low "unpredictability" as signs of AI, but these qualities can simply reflect a writer's personal style.

The consequences have already been severe. Last month, publisher Minotaur dropped a $2 million book deal over concerns that author Jerry Falade used AI—something he vehemently denies. Thierry Rignol, a French national, sued Yale after a professor used GPTZero to scan his final exam; Rignol failed and received a one-year suspension. The lawsuit argues that "AI surveillance and detection tools are known to unfairly target non-native English speakers." In February, an Adelphi University student won a lawsuit against the school after his professor claimed he used AI in an essay based on an AI detector scan. Outside academia, journalist and Verge contributor Kat Tenbarge was accused of using AI by Jack Osbourne (Ozzy Osbourne's son) in a broadcast to over 3.5 million followers, with Osbourne flaunting Getsolved detector results as "proof." Tenbarge refuted the claim, but Osbourne has not retracted the video, leaving her to deal with trolls. Even Turnitin has cautioned that its tool "may not always be accurate" and "shouldn't be used to take actions against a student." Grammarly warns that users "should never rely on the results of an AI detector alone." GPTZero states "no AI detector can ever truly be 100% perfect." OpenAI shut down its own AI writing detector in 2023 due to low accuracy.

Recognizing the unreliability, leading institutions have backed away. Yale University, Johns Hopkins University, Vanderbilt University, Georgetown University, and others have disabled or restricted AI detection tools. MIT warns that "AI detectors don't work." In their place, educators are adopting human-centered approaches: the University of Chicago suggests slowing students' reading, breaking up long writing assignments, and requiring reflection on work. Stanford advises holding assessments in classrooms. MIT encourages professors to leave room for students to disclose AI use without penalty. Meanwhile, platforms are moving in the opposite direction—Substack has integrated Pangram into its app to scan blogs for suspected AI content, and LinkedIn added a "seems like AI slop" button. The result is an era of pervasive distrust, where readers constantly question authenticity and writers work to avoid sounding like machines. Organizations like the Authors Guild are fighting back by offering "Human Authored" certifications, and platforms like Wikipedia have banned AI-generated articles outright.

Context & Analysis

The rise of AI detectors reflects genuine educator anxiety about generative AI in classrooms and professional writing, but it has outpaced the tools' actual capability. These detectors do not work like traditional plagiarism checkers—they do not match text against a database of known sources but instead rely on algorithmic guesses about writing patterns. This makes them fundamentally less verifiable: an accusation based on a detector scan cannot be backed up by a concrete match the way a plagiarism flag can. The body of evidence within the article shows consistent failure on non-native English speakers, yet these populations remain unprotected; vendors continue to claim accuracy despite disclaimers in their own terms of service.

The damage inflicted by false accusations has become tangible and public. When a high-profile case like publisher Minotaur's $2 million book deal cancellation or a student's year-long suspension goes public, others follow—either validating the use of detectors or emboldening further accusations. Platforms amplifying detection tools (Substack, LinkedIn) are institutionalizing the suspicion. Meanwhile, the institutions with the most authority—Yale, MIT, Johns Hopkins—have opted out, signaling to educators that the tools are unreliable. The body notes that OpenAI itself abandoned its detector, a striking vote of no-confidence from the very company at the center of the AI anxiety.

FAQ

How accurate are these AI detectors?
The detectors have high false positive rates despite vendor claims. A 2023 Stanford study found they falsely flagged essays by non-native English speakers more often than native speakers. OpenAI shut down its own detector in 2023 due to low accuracy, and even vendors acknowledge limitations—Grammarly warns users "should never rely on the results of an AI detector alone," while GPTZero says "no AI detector can ever truly be 100% perfect."
What have schools done in response?
Yale University, Johns Hopkins University, Vanderbilt University, Georgetown University, and MIT have disabled or restricted AI detection tools. Instead, schools like Stanford and the University of Chicago are encouraging professors to hold in-classroom assessments, ask students to reflect on their work, and allow students to disclose AI use without penalty.
Are there examples of people wrongly accused?
Yes. A Yale student sued the university after failing an exam based on a GPTZero scan; an Adelphi University student won a lawsuit against the school after being accused of using AI in an essay. Publisher Minotaur dropped a $2 million book deal over AI-use concerns, which the author Jerry Falade vehemently denies.

Get AI news like this every morning

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytime

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Next articleStrong earnings ease AI-concentration worries in U.S. stock rally

The AI news that matters, in one minute each morning.

Sign up free