
According to Bradley Emi, CTO of Pangram (an AI text detector), post-training safety guardrails in ChatGPT, Claude, and Gemini cause their outputs to exhibit "mode collapse"—where models fixate on narrow phrasing patterns instead of matching human linguistic variety.
This constraint makes their text statistically detectable.
By contrast, base models and specialized fine-tunes that lack such guardrails maintain broader expressive range and evade detection, though watermarks are expected to remain effective across all model types.
What happened
Bradley Emi, CTO of AI text detector Pangram, argues that safety guardrails in systems like ChatGPT, Claude, and Gemini narrow their expressive range through "mode collapse," making their outputs detectable—whereas base models (raw models before post-training) and specialized fine-tunes write with more variety and avoid detection.
Why it matters
The finding suggests that the safety measures deployed in mainstream AI systems create a trade-off: they constrain the model's behavioral range to avoid dangerous or censored outputs, but that constraint leaves a statistical fingerprint that detection tools can exploit. For businesses deploying or auditing AI-generated text, this implies that safety-aligned models may be easier to identify as machine-written, while organizations using base models or fine-tunes for narrower domains face higher detection risk.
What to watch
Watermarks are noted as a detection method likely to work reliably across model variants, including base models with higher stylistic variety, suggesting watermarking may become the more durable safeguard as detection tools evolve.
Ask the AI about this article →
The core insight from Emi's research is that the safety measures most widely deployed in consumer-facing AI systems—the guardrails that prevent harmful, biased, or censored outputs—come at a cost to linguistic diversity. Rather than sampling from the full probability distribution of human language, these models concentrate their output probability on a narrower, safer range of phrasings. This concentration is what enables detection: statistical tools can identify the characteristic fingerprint of that constrained distribution.
The implication cuts two ways. First, it shows a real tension between safety and authenticity: the guardrails that make systems responsible and compliant also make them identifiable. Second, it hints at an asymmetry in the AI landscape—base models and domain-specific fine-tunes, which lack post-training safety guardrails, maintain a broader stylistic range and thus evade standard detection methods. The caveat about watermarks is significant: Emi notes that watermarking appears orthogonal to this guardrail effect, suggesting it may be a more robust detection approach as the cat-and-mouse game between generative models and detectors evolves.
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytime
Ask AI anything about this article. Q&As are published on this page for other readers too.
As AI technology matures, the bottleneck in the industry is moving beyond semiconductor constraints like GPUs…

OpenAI has launched an Apple Messages plug-in for ChatGPT that lets users connect their Messages inbox to the…

Amazon Bedrock now supports OpenAI GPT-5.6 models (Sol, Terra, and Luna variants) across more than 25 AWS Regi…

Slack introduced Slack Code, a new feature that lets teams collaborate with AI coding agents (Claude, Devin, G…

Cisco is transforming its digital customer experience (DCX) strategy by embedding AI throughout customer journ…

Mastercard CEO Michael Miebach introduced "Agent Pay" last April, a payment framework that allows AI agents to…
