AIToday
AI Safety & AlignmentAI Regulation & PolicyAI Business & IndustryLatent SpacePublished: Sep 15, 2026, 16:00 JST

Dario pledges embedded evaluators as AEF-1 standard lands

Dario pledges embedded evaluators as AEF-1 standard lands

3 Key Points

  1. What happened

    Anthropic CEO Dario Amodei published a rare personal blog post committing Anthropic unilaterally to embedded third-party evaluators — giving them desks, access badges, and company laptops in its offices.

  2. Why it matters

    The AI Evaluator Forum also published AEF-1, a proposed baseline for independent evaluations covering access, conflicts of interest, funding, recusal, and transparency — shifting safety from principles toward verifiable process.

  3. What to watch

    The test is whether evaluators recruited by big labs can genuinely be independent; Kevin Bass posted the highest-engagement thread alleging financial entanglement between Anthropic and parts of the safety ecosystem.

WHO IT HITSThis lands on AI safety and evaluation teams at frontier labs, who would gain employee-like access to internal risk workflows, and on third-party auditors like METR that would be recruited to verify safety commitments.

Ask the AI about this article →

Summaries like this, in your inbox every morning.

Context & Analysis

The pacing debate returned this week after first emerging in July with Pacing the Frontier. Dario Amodei, lead author on the original, wrote a rare personal blog post spelling out how he sees pacing play out, and the AI Evaluator Forum — formed in December 2025 — happened to publish AEF-1, its expectations for what third-party evaluators should do, at the same time. The forum's members are presumably the leading third-party auditors that big labs would recruit for self-regulation.

The proposal has three parts: Embedded Evaluators, Democratic Coordination, and Global Coordination. Amodei's promise of desks, badges, laptops, and access to workspaces and tools mostly comparable to internal risk assessment teams goes further than standard domestic self-regulation playbooks. The AEF-1 baseline covers access, conflicts of interest, funding relationships, recusal, and transparency — the same issues raised by critics.

A sharp public split runs underneath the announcement. Bilal Chughtai left Google DeepMind and called for pacing and more transparency; Dan Selsam argued models may appear aligned under evaluation while hiding misalignment. On the other side, Shashank/Sayash Kapoor and Lennart Heim frame rogue agent incidents as a security and governance problem, and Aidan Gomez argued against a few Silicon Valley companies becoming AI gatekeepers. The test is likely whether embedded evaluators can be genuinely independent in practice, and whether Kevin Bass's allegations of financial entanglement between Anthropic and parts of the safety ecosystem hold up.

FAQ
What exactly is Anthropic committing to?
Dario Amodei's post says Anthropic is unilaterally committing now to embedded third-party evaluators with desks in its offices, access badges, company laptops, and workspace permissions mostly comparable to internal risk assessment teams.
What is AEF-1?
AEF-1 is a proposed baseline for independent third-party AI evaluations published by the AI Evaluator Forum, covering access, conflicts of interest, funding relationships, recusal, and transparency.
What are the three steps of Amodei's pacing framework?
The post lays out Embedded Evaluators, Democratic Coordination among frontier AI companies in democratic countries, and Global Coordination with authoritarian governments to the extent possible.

Get the latest AI Safety & Alignment news every morning

For example, today's edition would include:

  • AIUC raises $40 million to audit frontier AI modelsSiliconANGLE AI · 1h ago
  • Amodei, Altman slowdown calls hit AI stocks, but fears called overblownYahoo Finance AI · 1h ago
  • Musk pitches mutual AI model checks ahead of Trump-Xi talksSemafor Tech · 1h ago

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleGPT-6 Astra shows belief-propagation-like inference without CoT