AIToday
Large Language ModelsAI Business & IndustryThe Rundown AIPublished: Sep 26, 2026, 01:00 JST

Anthropic: Claude leads 26% of measured AI research

Anthropic: Claude leads 26% of measured AI research

3 Key Points

  1. What happened

    Anthropic published internal estimates showing Claude at the "leads" level for 26% of its measured AI research and development work in August, up from under 1% in February, with humans supervising.

  2. Why it matters

    This exposes how far automation has progressed inside a company that has publicly called for pacing AI progress, making that call for restraint more concrete.

  3. What to watch

    The ratings leave room for disagreement about where work falls on the scale, and Claude's scores matched human ratings exactly 59% of the time.

WHO IT HITSAI safety teams and research evaluators will need to judge whether human review can keep up with more experiments, given that full autonomy was zero and about 30,000 agents ran at once.

Not sure about something? Ask the AI

Questions and answers are published on this page.

Summaries like this, in your inbox every morning.

Context & Analysis

Anthropic's internal estimates arrive after CEO Dario Amodei called for pacing AI progress earlier this month and warned on September 12 that AI helping build its successors could speed development beyond people's ability to understand and control the resulting systems. His immediate commitment was to bring outside evaluators into Anthropic with access similar to employees, while broader limits and shared standards would depend on coordination among companies and governments. The measurements make that call for restraint more concrete by exposing how far automation has progressed inside the company asking others to slow the pace.

Anthropic reported about 30,000 agents running at once on its largest internal platform, with roughly one in 47,000 decisions blocked across more than one billion decisions in August. Safety work received about 6% of R&D computing resources during July 13–20, a snapshot that excluded safeguards classifiers. Anthropic's August risk report, covering conditions as of July 15, said the company had yet to evaluate its automated offline monitoring from start to finish, leaving open how many problems went undetected.

Public transparency depends on who can test the numbers. Amodei acknowledges that companies choose what their own disclosures include and omit, and evaluators would need to inspect underlying records, challenge classifications and report adverse findings for outsiders to judge whether the measures hold up. Whether endorsements from Altman and Musk produce comparable access and public reporting across labs remains an open question, so the stakes may hinge on the access and reporting rules still unsettled in Anthropic's Accenture partnership.

FAQ
How did Anthropic measure Claude's role in its research?
Anthropic applied Epoch AI's scale to a fixed set of tasks, weighted by the human time they take. Claude itself rated the tasks, and its scores matched human ratings exactly 59% of the time.
What does "leads" mean in this context?
"Leads" means Claude completes most of a task under human supervision. It reached "collaborates or above" for over 90% of the measured work, while full autonomy was zero.
What did Anthropic announce with Accenture?
On September 18, Anthropic announced an evaluation partnership with Accenture, led by its Faculty business, covering red teaming, alignment assessments and safeguards testing. Anthropic will fund the work directly, and access and reporting rules were left unsettled.
The Rundown AIRead Original Article

Get the latest Large Language Models news every morning

For example, today's edition would include:

  • Warp Agent takes on HR chores, CEO eyes post-Workday worldSiliconANGLE AI · 1h ago
  • OpenAI agents in rogue swarm hack; model launches roll onSiliconANGLE AI · 1h ago
  • Davis: AI agents need verifiable computing, not trustSiliconANGLE AI · 1h ago

AI-summarized, only the topics you pick: one digest a day via Email, LINE, or Slack.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.

Questions and answers are published on this page.

Related Articles

Next articleHarker: AI cut the entry-level training subsidy