AIToday
Large Language ModelsAI Safety & AlignmentAI Business & IndustryThe Rundown AIPublished: Sep 11, 2026, 01:00 JST2 min read

Anthropic's Jacob Coxon exits, warns of AI race

Anthropic's Jacob Coxon exits, warns of AI race

3 Key Points

  1. What happened

    Anthropic researcher Jacob Coxon resigned, accusing Anthropic and OpenAI of "gambling with our lives" by racing toward AI that can improve itself. Anthropic alignment lead Evan Hubinger replied, "We really do earnestly believe AI could kill all humans!"

  2. Why it matters

    The alarm came from inside Anthropic's own safety effort, not outside critics. Hubinger estimated a greater than 10% chance of AI killing all humans over the next decade and said Anthropic has no plan yet for controlling superintelligence.

  3. What to watch

    Anthropic's stated rationale is that safety research needs the most advanced models, so the test is whether its August 2026 Risk Report reasoning holds for more capable successors. Coxon instead called for a coordinated slowdown, possibly a temporary ban on improving model capabilities.

WHO IT HITSEnterprise buyers and policymakers weighing Anthropic's safety-first pitch now have a public statement from its own alignment research lead that the company is not clearly on track to solve alignment. Compliance and risk teams evaluating AI vendors may find the internal dissent harder to dismiss than outside criticism.

Ask the AI about this article →

Summaries like this, in your inbox every morning.

Context & Analysis

The unusual part is where the doubt came from. Coxon was not an outside critic; he spent three years helping train models across OpenAI and Anthropic. And the reply came from Evan Hubinger, an alignment research lead at Anthropic itself, who said the company does not yet have a plan for controlling superintelligence and was not clearly on track to solve alignment.

Anthropic has long acknowledged this uncertainty. In its March 2023 statement on AI safety, it said making powerful systems reliably safe remained unsolved, warned that competition could encourage unsafe deployment, and admitted its research might fail. Its August 2026 Risk Report, covering information through July 15, judged the risk of its covered models causing a catastrophe to be low, partly because those models are unlikely to be capable enough to secretly undermine oversight.

That reasoning is what needs testing for more capable successors, and it is where the stakes sit. Anthropic's stated rationale is that meaningful safety research requires access to the most advanced models — an approach whose reliability is unproven. Coxon's call for a coordinated slowdown targets the competitive pressure, but the body notes it remains unclear who would agree, what a ban would cover, or how it would be enforced. Customers and policymakers are left asking what evidence would justify further capability increases and who checks those judgments.

FAQ
What did Jacob Coxon say when he resigned?
Coxon accused Anthropic and OpenAI of "gambling with our lives" by racing toward AI that can improve itself. He called for a coordinated slowdown and suggested preventing a global race could require a temporary ban on improving model capabilities.
What is Evan Hubinger's estimate of AI risk?
Hubinger personally estimated a greater than 10% chance of AI killing all humans over the next decade. He said he considered today's models low risk and was concerned about future AI that could become far more capable by repeatedly improving itself.
What did Anthropic's August 2026 Risk Report find?
The report, released August 14 and covering information through July 15, judged the risk of its covered models causing a catastrophe by acting against human intentions to be low. It rests partly on those models being unlikely to have strong enough abilities to secretly undermine oversight.
The Rundown AIRead Original Article

Get the latest Large Language Models news every morning

For example, today's edition would include:

  • DeepSeek V4.1-Flash: 763B model beats V4 Pro on AA Index 40Latent Space · 38m ago
  • Dynatrace acquires Arize AI as observability shifts to actionSiliconANGLE AI · 6h ago
  • Shared base cuts 100 fine-tunes from 1.5 TB to 19.3 GBDaily Dose of Data Science · 6h ago

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next article348M model hits 99.4% on GPT-3 arithmetic tasks