
Researchers released WeirdChat, a public catalog of over 175,000 annotated transcripts documenting over 1,300 behavioral patterns discovered in frontier open-weight language models. The behaviors range from benign (making up a user's name) to dangerous (encouraging self-harm), surfaced using automated techniques designed to elicit unexpected outputs in simulation.
Summaries like this, in your inbox every morning.
For example, today's edition would include:
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.
Co-leads of the UN Global Dialogue on AI Governance in Geneva, where 170 countries met in July, are proposing…

Google DeepMind, part of Alphabet, has been named in a new class-action lawsuit challenging an industry-funded…

After a cabinet-level meeting in New York with Chinese Vice Premier He Lifeng and others, US Treasury Secretar…

A Reddit user, /u/gareth789, who says they work on an AI project, floated the idea of an AI twin built around…

Reddit user /u/AkindaGood_programer posted that using LLMs to find gaps in their knowledge by having the model…

Anthropic published a threat intelligence report on how advanced AI systems are used under heavy real-world wo…
