
What happened
Anthropic published an index scoring its own development work, and says 26 percent of it now sits at AL4 — Epoch AI's label for "AI leads" — up from under one percent in February.
Why it matters
At AL4 the work is still handed over by a person and shipped by a person, so the 26 percent figure counts human hours, not decisions taken away from engineers.
What to watch
Whether the number holds up hinges on the scoring, which Anthropic says its own model produced, with official scores matching human judgment 59 percent of the time.
WHO IT HITSAnthropic's own engineers and researchers are the ones whose work is being scored, and the people reading the report — including regulators and rival labs — now have a company-published number for how much of AI research a model already runs.
Summaries like this, in your inbox every morning.
The report lands alongside a call from Anthropic CEO Dario Amodei to slow development at the AI frontier in a coordinated way. Anthropic argues that the public needs more insight into how models are built, and the metrics are meant to complement capability tests that measure what models can do. The central measure sits on a scale from Epoch AI running from AL0 (no AI) to AL5 (fully autonomous), and Claude hits AL5 nowhere. More than 90 percent of the work reaches at least AL3.
The harder question is what the 26 percent counts. Anthropic listed how employees spent their work time in July and scored each activity, with tasks that eat up a lot of time counting for more, so the figure reflects the human hours a task takes — not how many decisions a model makes or whether it has a say in research direction. Anthropic's own cross-check suggests the boundary between "collaborates" and "leads" is slippery: when two employees rated the same work area, they agreed on the level only about a third of the time, and one level is exactly the jump from AL3 to AL4.
A second set of figures covers oversight. On Anthropic's most-used internal platform, about 30,000 agents run at the same time, with a real-time monitor meant to block things like copying model weights; out of more than a billion decisions in August, it stopped 0.002 percent. A second monitor flags about 100,000 logs a week after the fact, and roughly 50 reach humans. Anthropic says it has only run this monitoring for a few months and is not sure it catches every behavior pattern. How much weight outsiders give the headline number may hinge on whether Anthropic's self-scoring is treated as a performance claim, as an argument for slowing down, or as both.
For example, today's edition would include:
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.
Sam Altman and Elon Musk backed Dario Amodei's call for a slowdown in model releases, while Marc Benioff at Dr…
Eva Brucherseifer and Jan Muehlig will present 'What would it take?

Meta Platforms shares are up 24.34% over the past month, as Muse became the #1 app in the App Store one week a…

Broadcom, Meta Platforms, and Microsoft are all trading around breakeven for the year

In the first posts from the new DeepMind Institute, researchers Rohin Shah and Anca Dragan argue visible chain…

OpenAI launched Astra for Law, pairing GPT-6 Astra with a legal search index covering over 230 million URLs of…
