Computer scientists published a technique called IRM (Implicit Reward Model) that identifies text generated by large language models (AI systems that write text) by analyzing patterns in publicly available AI models. Unlike previous detection methods, IRM requires no labeled examples, no preference data collection, and no retraining—it works immediately on models already in use.
IRM reuses the built-in 'decision-making logic' (reward models) that AI labs already embed in instruction-tuned models to make them safer and more helpful. Rather than asking "does this text match human writing?", it asks "would an AI trained to be helpful rank this text highly?" The difference: previous detection approaches needed custom training for each new AI system; IRM adapts to new models automatically.
For schools and businesses, this matters because detecting AI-written homework, job applications, and customer reviews becomes cheaper and faster—no need to license expensive third-party tools or wait for custom detection to be built. For content platforms (news sites, forums, academic publishers), instant detection without setup means they can flag AI content in real time, reducing fake content sneaking through moderation.
Ask the AI about this article →
For example, today's edition would include:
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.
Anthropic announced Claude Fable 5.1 and Claude Mythos 5.1 on September 1

A technical explainer compares three LLM serving strategies—static, dynamic, and continuous batching

Anthropic's latest model, Claude Fable 5.1, is now available on Snowflake Cortex AI

The Allen Institute for AI released BenchMIRT, a method to audit AI benchmarks question-by-question

Google has reportedly approached major studios like Disney, Warner Bros

OpenAI shared new details on its forthcoming Astra model, which the company says is the first large language m…
