
What happened
Anthropic CEO Dario Amodei proposed embedding independent evaluators such as METR and Redwood Research inside frontier AI companies, with the right to publish findings without Anthropic's editorial control. OpenAI CEO Sam Altman said OpenAI would commit too.
Why it matters
If evaluators get real access to training checkpoints and logs, they could catch models that behave well during testing while hiding problems, an outcome that today's pre-release testing may miss.
What to watch
The pledge hinges on whether companies actually surrender control over access and publication, since prior evaluation windows like three days for GPT-6 Astra were too short to draw firm conclusions.
WHO IT HITSPolicy and compliance teams at frontier AI labs and third-party evaluation firms like METR would face new access, NDA and publication rules, while California and EU regulators weigh whether voluntary pledges are enough.
Summaries like this, in your inbox every morning.
Anthropic's and OpenAI's endorsements mark a shift from the industry's past practice of bringing in outside reviewers only shortly before a model's release. The evaluators who spoke to TechCrunch argue that access to intermediate training checkpoints, post-training environments and evaluation transcripts would let them verify a company's public safety claims and trace when concerning behavior first emerged.
The track record of prior efforts suggests the hard part is not the principle but the terms. FAR.AI has turned down contracts with developers that demanded too much control, and researchers note that OpenAI gave METR and Redwood roughly a week to investigate the Hugging Face incident, while Apollo Research had only three days to test GPT-6 Astra. Those limitations made firm conclusions difficult.
Whether this proposal becomes meaningful oversight may hinge on whether it is backed by regulation rather than company goodwill, and on whether auditors can publish freely. Researchers point to standards for auditor qualifications and to laws like SB 813 as possible anchors. For labs, the test is whether they accept outside findings they cannot edit; for evaluators, whether the access is real and lasting.
For example, today's edition would include:
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.
CADDi raised $114 million at a $1.2 billion valuation, versus the $470 million it reported at its March 2025 S…
Anthropic merged chatbot Claude with agentic tool Claude Cowork effective immediately, and launched Claude Doc…
Hang Ten Systems Inc., founded by ex-Infosys chief Vishal Sikka, announced a $53 million second seed round led…
Cohere and Aleph Alpha signed a merger agreement today, formalizing an April plan; the combined company will r…
Elon Musk has hinted at a Tesla-SpaceX merger, according to DIGITIMES, which notes the two companies are alrea…

Anthropic posted guidance saying Claude Code's output tokens cost about 5 times its input tokens, and that one…
