AIToday
Large Language ModelsAI Safety & AlignmentLessWrong AIPublished: Jul 22, 2026, 13:01 JST

WeirdChat: 175,000 AI model behaviors cataloged automatically

WeirdChat: 175,000 AI model behaviors cataloged automatically

Researchers released WeirdChat, a public catalog of over 175,000 annotated transcripts documenting over 1,300 behavioral patterns discovered in frontier open-weight language models. The behaviors range from benign (making up a user's name) to dangerous (encouraging self-harm), surfaced using automated techniques designed to elicit unexpected outputs in simulation.

Not sure about something? Ask the AI

Summaries like this, in your inbox every morning.

Get the latest Large Language Models news every morning

For example, today's edition would include:

  • AI twin idea: help readers or replace them?r/artificial · 6h ago
  • /u/AkindaGood_programer: LLMs work well for finding knowledge gapsr/artificial · 6h ago
  • Anthropic report: server choice hinges on token tasksDIGITIMES Asia · 9h ago

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleAnthropic, OpenAI in robotics race; acquisition talks with Physical Intelligence surface