AIToday
Large Language ModelsAI Coding AssistantsAI Business & IndustryTHE DECODERPublished: Sep 19, 2026, 01:00 JST

Anthropic: 26 percent of dev work at AL4, a quarter of research

Anthropic: 26 percent of dev work at AL4, a quarter of research

3 Key Points

  1. What happened

    Anthropic published an index scoring its own development work, and says 26 percent of it now sits at AL4 — Epoch AI's label for "AI leads" — up from under one percent in February.

  2. Why it matters

    At AL4 the work is still handed over by a person and shipped by a person, so the 26 percent figure counts human hours, not decisions taken away from engineers.

  3. What to watch

    Whether the number holds up hinges on the scoring, which Anthropic says its own model produced, with official scores matching human judgment 59 percent of the time.

WHO IT HITSAnthropic's own engineers and researchers are the ones whose work is being scored, and the people reading the report — including regulators and rival labs — now have a company-published number for how much of AI research a model already runs.

Not sure about something? Ask the AI

Summaries like this, in your inbox every morning.

Context & Analysis

The report lands alongside a call from Anthropic CEO Dario Amodei to slow development at the AI frontier in a coordinated way. Anthropic argues that the public needs more insight into how models are built, and the metrics are meant to complement capability tests that measure what models can do. The central measure sits on a scale from Epoch AI running from AL0 (no AI) to AL5 (fully autonomous), and Claude hits AL5 nowhere. More than 90 percent of the work reaches at least AL3.

The harder question is what the 26 percent counts. Anthropic listed how employees spent their work time in July and scored each activity, with tasks that eat up a lot of time counting for more, so the figure reflects the human hours a task takes — not how many decisions a model makes or whether it has a say in research direction. Anthropic's own cross-check suggests the boundary between "collaborates" and "leads" is slippery: when two employees rated the same work area, they agreed on the level only about a third of the time, and one level is exactly the jump from AL3 to AL4.

A second set of figures covers oversight. On Anthropic's most-used internal platform, about 30,000 agents run at the same time, with a real-time monitor meant to block things like copying model weights; out of more than a billion decisions in August, it stopped 0.002 percent. A second monitor flags about 100,000 logs a week after the fact, and roughly 50 reach humans. Anthropic says it has only run this monitoring for a few months and is not sure it catches every behavior pattern. How much weight outsiders give the headline number may hinge on whether Anthropic's self-scoring is treated as a performance claim, as an argument for slowing down, or as both.

FAQ
What does AL4 actually mean at Anthropic?
Epoch AI's scale runs from AL0 (no AI) to AL5 (fully autonomous), and AL4 is labeled "AI leads." In Anthropic's example, an engineer hands Claude a bug report, and Claude analyzes, fixes, and tests it without questions — but a human still decides whether it ships.
Who judged whether the work counted as AL4?
Claude did. Agents gathered evidence from Slack and internal documents, and another Claude model assigned the levels, which Anthropic admits could repeat the same mistakes as the system it is checking.
How much of Anthropic's compute goes to safety work?
About six percent of the compute for AI research, in a sample week in July. Anthropic says that is low partly because safety work is mostly experiment design, which costs researcher time rather than chips.

Get the latest Large Language Models news every morning

For example, today's edition would include:

  • Sam Altman, Elon Musk back Amodei's AI slowdown callSiliconANGLE AI · 1h ago
  • KDE at 30: Kadai AI-native desktop plan splits AkademyThe Register (AI/ML) · 1h ago
  • Meta rebounds 24.34% as Muse hits #1 in App StoreYahoo Finance AI · 1h ago

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleGitHub Podcast: 5 AI hot takes don't hold up