AIToday
Large Language ModelsAI Business & IndustryHacker NewsPublished: Sep 18, 2026, 06:00 JST

Anthropic: Claude "leads" 26% of its AI R&D

Anthropic: Claude "leads" 26% of its AI R&D

3 Key Points

  1. What happened

    Anthropic's prototype index rates automation from AL0 to AL5; as of August 2026 Claude is not fully autonomous on any measured AI R&D subset, but 'leads' 26%.

  2. Why it matters

    A frontier lab is quantifying how much AI builds AI, which the body says helps connect inputs like compute to outputs like capabilities and inform any pacing effort.

  3. What to watch

    The numbers are conservative, depend on a Claude judge, and cover one week of compute; watch whether independent third-party evaluators Anthropic plans to embed can verify them.

WHO IT HITSPolicymakers and regulators weighing rules on frontier AI development gain a concrete template for what labs could disclose, while rival lab safety and compliance teams may face pressure to publish comparable metrics.

Not sure about something? Ask the AI

Summaries like this, in your inbox every morning.

Context & Analysis

Anthropic says AI systems are becoming exponentially more powerful and have begun to automate more of the process of building themselves. As the world considers slowing the pace of frontier AI development, the company argues the public needs more information. That framing sets up the piece not as a capability boast but as a transparency exercise: it offers measurement tools for how much AI is building its next version, how well Anthropic can oversee and intervene in agent actions, and where compute goes.

The three measurements connect to Anthropic's broader policy proposals. Its Responsible Scaling Policy risk reports already publish capability evaluations, while the Advanced AI Framework proposes transparency obligations governments could require. The measurements here focus on the production process, so they can be correlated with those capability evaluations. The body also notes Anthropic plans to embed independent third-party evaluators from multiple organizations with access comparable to internal risk teams, and says external red-teaming has happened before, with METR independently testing the offline monitoring platform.

The stakes hinge on verifiability and comparability. Anthropic acknowledges obstacles to cross-lab comparison, including the lack of a common methodology and its use of its own models as judges. The compute snapshot also covers one week, which the body says shows the measurement can be made but not a meaningful trend. How much outside parties can trust these numbers, and whether other developers adopt similar reporting, is likely to determine their influence.

FAQ
How autonomous is Claude at Anthropic?
As of August 2026, Claude does not operate fully autonomously on any measured subset of AI R&D work, but it 'leads' 26% of that work and over 90% is at or above 'AI collaborates.'
How much of Anthropic's compute goes to safety?
Over one examined week, about 6% of compute that went to AI R&D was allocated toward safety, and about 12% of compute for AI-driven AI R&D went to safety.
How many agents does Anthropic monitor?
As of August 2026, there were approximately 30,000 agents doing research and engineering work at Anthropic at any one time on its most-used internal platform.

Get the latest Large Language Models news every morning

For example, today's edition would include:

  • ByteDance's AI agent phone hits app wallDIGITIMES Asia · 3h ago
  • OpenAI launches Astra for Law, a GPT-6 setup for legal researchSiliconANGLE AI · 6h ago
  • Google opens CC to families of six as shared AI agentSiliconANGLE AI · 6h ago

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleplzdontkillus '21M+ AI risk views' shrinks to ~2M, fellow says