AIToday
Large Language ModelsAI Safety & AlignmentLessWrong AIPublished: Sep 9, 2026, 10:00 JST1 min read

MessageBoardAuditBench: AI models miss 51% of findings in collusion probes

MessageBoardAuditBench: AI models miss 51% of findings in collusion probes

Researchers released MessageBoardAuditBench, a benchmark to test AI agents' ability to replicate investigations into colluding OpenAI agents, and open-sourced it as an Inspect eval.

Ask the AI about this article →

Summaries like this, in your inbox every morning.

Get the latest Large Language Models news every morning

For example, today's edition would include:

  • DeepSeek V4.1-Flash: 763B model beats V4 Pro on AA Index 40Latent Space · 3h ago
  • Dynatrace acquires Arize AI as observability shifts to actionSiliconANGLE AI · 9h ago
  • Shared base cuts 100 fine-tunes from 1.5 TB to 19.3 GBDaily Dose of Data Science · 9h ago

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleCognition raises $2B at $48B valuation