
What happened
Anthropic researcher Jacob Coxon resigned, accusing Anthropic and OpenAI of "gambling with our lives" by racing toward AI that can improve itself. Anthropic alignment lead Evan Hubinger replied, "We really do earnestly believe AI could kill all humans!"
Why it matters
The alarm came from inside Anthropic's own safety effort, not outside critics. Hubinger estimated a greater than 10% chance of AI killing all humans over the next decade and said Anthropic has no plan yet for controlling superintelligence.
What to watch
Anthropic's stated rationale is that safety research needs the most advanced models, so the test is whether its August 2026 Risk Report reasoning holds for more capable successors. Coxon instead called for a coordinated slowdown, possibly a temporary ban on improving model capabilities.
WHO IT HITSEnterprise buyers and policymakers weighing Anthropic's safety-first pitch now have a public statement from its own alignment research lead that the company is not clearly on track to solve alignment. Compliance and risk teams evaluating AI vendors may find the internal dissent harder to dismiss than outside criticism.
Ask the AI about this article →
Summaries like this, in your inbox every morning.
The unusual part is where the doubt came from. Coxon was not an outside critic; he spent three years helping train models across OpenAI and Anthropic. And the reply came from Evan Hubinger, an alignment research lead at Anthropic itself, who said the company does not yet have a plan for controlling superintelligence and was not clearly on track to solve alignment.
Anthropic has long acknowledged this uncertainty. In its March 2023 statement on AI safety, it said making powerful systems reliably safe remained unsolved, warned that competition could encourage unsafe deployment, and admitted its research might fail. Its August 2026 Risk Report, covering information through July 15, judged the risk of its covered models causing a catastrophe to be low, partly because those models are unlikely to be capable enough to secretly undermine oversight.
That reasoning is what needs testing for more capable successors, and it is where the stakes sit. Anthropic's stated rationale is that meaningful safety research requires access to the most advanced models — an approach whose reliability is unproven. Coxon's call for a coordinated slowdown targets the competitive pressure, but the body notes it remains unclear who would agree, what a ban would cover, or how it would be enforced. Customers and policymakers are left asking what evidence would justify further capability increases and who checks those judgments.
For example, today's edition would include:
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.
Michael Burry, famous for The Big Short, is short Nvidia, Palantir and Tesla, and in his Substack newsletter s…

Investors have three creative routes to Anthropic exposure before its expected IPO: buying Alphabet, Amazon, o…

DeepSeek launched V4.1-Flash, a 763B-parameter open-weight model with a causal encoder-decoder architecture

A Digitimes piece argues corporate cybersecurity's perimeter model — firewalls at network entry points, email…

Much of the attention on AI infrastructure buildouts is now tied to sheer compute power, with dominance define…

Barron's reported September 10 that Kepler Computing emerged from stealth with a memory architecture using fer…
