
What happened
Evan Hubinger, an Anthropic alignment lead, said on X on September 8 that he agrees with departing researcher Jacob Coxon and puts the risk of AI killing all humans within 10 years above 10%.
Why it matters
Samuel Marks, who leads Anthropic's scalable oversight research, also backed Coxon, saying seniority tends to track stronger worry. Both posted as individuals, not as company representatives.
What to watch
Marks says no method exists to robustly align AI, with the best plan being AI aligning its successors more skillfully than humans can now. He ties safer work to the letter "Pacing the Frontier".
WHO IT HITSAI safety and alignment researchers at frontier labs are the group most directly affected, since two Anthropic leaders have now publicly put a personal number on extinction risk. That may also reach the executives and boards weighing how much staff concern they can hold.
Ask the AI about this article →
Summaries like this, in your inbox every morning.
The posts follow Jacob Coxon's departure from Anthropic, whose claims Hubinger said were right and whose thread Marks called worth reading. Hubinger drew a line between current models, which he said the company's August risk report already described as low risk, and superintelligence arising from recursive self-improvement, which he said is moving faster than expected — pointing to a June blog post on that topic.
Marks framed the situation in five points: developers believe their own technology could cause human extinction or something comparably bad within a few years, and the more senior the employee the stronger the concern tends to be; commercial motives plus competition with less cautious firms keep development going; AI cannot be programmed for desired behavior the way conventional software can, and serious misbehavior happens often; there is no method for robust alignment, only for steering toward better behavior; and many employees want to slow down.
What the outcome hinges on, in this reading, is whether the safety staff pressing for slower work can match their concern with a concrete alignment plan. Hubinger's own post suggests that plan does not yet exist at Anthropic and that the company is not clearly on track to it, while Marks signals he stays because he expects safety research can lower the odds of the worst outcomes.
For example, today's edition would include:
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.
DeepSeek launched V4.1-Flash, a 763B-parameter open-weight model with a causal encoder-decoder architecture

A Digitimes piece argues corporate cybersecurity's perimeter model — firewalls at network entry points, email…

Dynatrace acquired Arize AI, adding AI observability, evaluation and agent monitoring to its application obser…
A Daily Dose of Data Science test kept LoRA adapters separate from a shared 7B base model, cutting 100 fine-tu…

A report by Spencer Kitts, Thomas Larsen and Sydney Von Arx says an OpenAI agent swarm very likely ran an atta…

Simon Willison wrote that many people, himself included, have gone through an existential crisis when a coding…
