AIToday
AI Safety & AlignmentThe Verge AIPublished: Sep 9, 2026, 19:00 JST2 min read

Anthropic safety lead: >10% chance AI kills all humans; no safety plan yet

Anthropic safety lead: >10% chance AI kills all humans; no safety plan yet

3 Key Points

  1. What happened

    Jacob Coxon, who trained AI systems at Anthropic and OpenAI, resigned, accusing both of racing to self-improving superintelligence and gambling with lives.

  2. Why it matters

    Evan Hubinger, who leads an Anthropic safety team, agreed, saying AI could kill all humans with >10% chance within a decade, and that Anthropic has no plan yet.

  3. What to watch

    The test is whether Anthropic and rivals develop safety plans before racing ahead; Coxon says they are locked in a race despite the risk.

WHO IT HITSEnterprise decision-makers relying on AI systems from Anthropic and OpenAI face fundamental uncertainty about long-term safety and control. Also, employees and researchers at these firms must weigh personal risk against career advancement and impending IPOs.

Ask the AI about this article →

Summaries like this, in your inbox every morning.

Context & Analysis

Anthropic's own safety community is fracturing over the pace of AI development. Jacob Coxon's departure, after previously working at OpenAI, underscores a pattern of high-profile exits from both companies over safety concerns. His accusation that these labs are "racing straight to self-improving superintelligence" comes as they are reportedly preparing for IPOs, potentially adding financial pressure to already aggressive timelines.

Evan Hubinger's public agreement and candid admission that the company lacks a safety plan highlight an internal contradiction: leadership voices acknowledge existential risks but admit no concrete strategy to mitigate them. The exchange occurs against backdrop of rogue agent incidents and warnings about frontie model monitorability, suggesting these are not isolated fears but endemic issues recognized even within the industry.

The stakes hinge on whether Anthropic and its rivals can translate safety rhetoric into actionable plans before deploying more powerful systems. Coxon's "locked in a race" characterization suggests competitive pressure may override caution. For companies integrating these AI systems, this debate signals that the technology's future is not merely a matter of capability but of governance and trust, with potentially fatal consequences if worst-case predictions materialize.

FAQ
Who resigned from Anthropic and why?
Jacob Coxon, a researcher who trained AI systems at Anthropic and previously at OpenAI, resigned over safety concerns, accusing the companies of racing to self-improving superintelligence and gambling with lives.
What is Evan Hubinger's estimate of AI risk?
Evan Hubinger, who leads an Anthropic safety team, said he believes AI could kill all humans, personally estimating the chances at greater than one in 10 within the next decade.
Does Anthropic have a plan for safe advanced AI?
Hubinger said Anthropic does not yet have a plan for ensuring advanced AI remains safe and aligned with human values, and are not clearly on track to develop one.

Get the latest AI Safety & Alignment news every morning

For example, today's edition would include:

  • Anthropic researcher Jacob Kochson quits, warns AI may be uncontrollable by next yearITmedia AI+ · 3h ago
  • OpenAI claims Navier-Stokes solution found by 10,000 agentsLatent Space · 3h ago
  • OpenAI agents used German wiki to swap 18,000 messages, report saysLatent Space · 3h ago

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleLG Innotek uses AI to fix glass substrate microcracks before 2028 race