AIToday
Large Language ModelsAI Safety & AlignmentAI Business & IndustryThe Verge AIPublished: Apr 24, 2026, 04:00 JST1 min read

Anthropic's restricted Claude Mythos AI leaked to unauthorized users before official launch, undermining its 'safety-first' brand

Anthropic's restricted Claude Mythos AI leaked to unauthorized users before official launch, undermining its 'safety-first' brand

3 Key Points

  1. Anthropic announced Claude Mythos, a new AI model it claimed was too capable at hacking and cybersecurity to release publicly, and began controlled testing with select companies. Within days of that announcement, Bloomberg reports unauthorized users gained access to the model anyway — a breach Anthropic is now investigating.

  2. The company built its reputation on prioritizing AI safety over speed, positioning itself as more cautious than competitors like OpenAI. This leak directly contradicts that core brand promise, showing a gap between Anthropic's stated security practices and what actually happened.

  3. For business leaders evaluating AI vendors, this signals risk: a company trusted with sensitive AI access failed to control that access from day one. For employees at companies in Anthropic's testing program, it raises questions about whether their proprietary data or use cases were exposed to the same unauthorized group.

Ask the AI about this article →

Get the latest Large Language Models news every morning

For example, today's edition would include:

  • Walmart settles opioid claims for $50MTop Companies AI · 49m ago
  • Tim Cook's legacy hinges on Apple's AI betTop Companies AI · 49m ago
  • AT&T Builds AI-First Legal CenterTop Companies AI · 49m ago

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleDelve, a startup handling AI security certifications, itself became the weak link in Context AI's data breach