AIToday
Large Language ModelsAI Safety & AlignmentITmedia AI+Published: Sep 10, 2026, 10:01 JST2 min read

Anthropic's Evan Hubinger puts 10-year AI extinction risk above 10%

Anthropic's Evan Hubinger puts 10-year AI extinction risk above 10%

3 Key Points

  1. What happened

    Evan Hubinger, an Anthropic alignment lead, said on X on September 8 that he agrees with departing researcher Jacob Coxon and puts the risk of AI killing all humans within 10 years above 10%.

  2. Why it matters

    Samuel Marks, who leads Anthropic's scalable oversight research, also backed Coxon, saying seniority tends to track stronger worry. Both posted as individuals, not as company representatives.

  3. What to watch

    Marks says no method exists to robustly align AI, with the best plan being AI aligning its successors more skillfully than humans can now. He ties safer work to the letter "Pacing the Frontier".

WHO IT HITSAI safety and alignment researchers at frontier labs are the group most directly affected, since two Anthropic leaders have now publicly put a personal number on extinction risk. That may also reach the executives and boards weighing how much staff concern they can hold.

Ask the AI about this article →

Summaries like this, in your inbox every morning.

Context & Analysis

The posts follow Jacob Coxon's departure from Anthropic, whose claims Hubinger said were right and whose thread Marks called worth reading. Hubinger drew a line between current models, which he said the company's August risk report already described as low risk, and superintelligence arising from recursive self-improvement, which he said is moving faster than expected — pointing to a June blog post on that topic.

Marks framed the situation in five points: developers believe their own technology could cause human extinction or something comparably bad within a few years, and the more senior the employee the stronger the concern tends to be; commercial motives plus competition with less cautious firms keep development going; AI cannot be programmed for desired behavior the way conventional software can, and serious misbehavior happens often; there is no method for robust alignment, only for steering toward better behavior; and many employees want to slow down.

What the outcome hinges on, in this reading, is whether the safety staff pressing for slower work can match their concern with a concrete alignment plan. Hubinger's own post suggests that plan does not yet exist at Anthropic and that the company is not clearly on track to it, while Marks signals he stays because he expects safety research can lower the odds of the worst outcomes.

FAQ
Who are the Anthropic researchers who agreed with Jacob Coxon?
Evan Hubinger, who leads an alignment research team, and Samuel Marks, who leads research on scalable oversight. Both said they were speaking as individuals, not for Anthropic.
What probability does Hubinger give for AI causing human extinction?
He said he personally estimates the probability at over 10% within the next 10 years. He also said the risk from current models is low.
What did Anthropic staff sign to support slowing down?
An open letter called "Pacing the Frontier", which Marks said he signed. He said many employees want to slow down to find safer ways to develop AI.

Get the latest Large Language Models news every morning

For example, today's edition would include:

  • DeepSeek V4.1-Flash: 763B model beats V4 Pro on AA Index 40Latent Space · 1h ago
  • Dynatrace acquires Arize AI as observability shifts to actionSiliconANGLE AI · 7h ago
  • Shared base cuts 100 fine-tunes from 1.5 TB to 19.3 GBDaily Dose of Data Science · 7h ago

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleNVIDIA, Australian partners plan 2-gigawatt DSX buildout by 2027