AIToday
Large Language ModelsHacker NewsPublished: Jul 10, 2026, 04:00 JST2 min read

Berkeley RDI launches AI agent benchmark across 55 industries with 1,500+ tasks

Key takeaway

  • Berkeley RDI has launched Agents' Last Exam, a large-scale benchmark to evaluate AI agents on real-world professional tasks across 55 industries.

  • The benchmark currently includes 1,500+ collected tasks from a 5,000-task target, with scores designed to be objective and comparable across domains.

  • This addresses a gap in how AI agents—autonomous systems that complete professional work—are measured in practical, economically valuable scenarios.

3 Key Points

  1. What happened

    Berkeley RDI, working with 300+ industry experts, has built Agents' Last Exam, a benchmark to measure AI agent performance on real-world, economically valuable professional tasks. The benchmark currently spans all 55 targeted sub-industries and has collected 1,500+ tasks toward a 5,000-task target, with scores designed to be objective and comparable across domains.

  2. Why it matters

    AI agents—systems that make decisions and complete work autonomously—are increasingly used in professional workflows, but there has been no standard way to measure how well they perform on practical, verifiable tasks. This benchmark creates that standard, making it possible for businesses to assess whether agents can reliably handle the long-horizon work that matters economically.

  3. What to watch

    The benchmark is aiming to reach 5,000 tasks total, covering most major fields of professional computer work. As more tasks are added and results accumulate, the benchmark will become a reference point for comparing agent capabilities across industries.

Ask the AI about this article →

Context & Analysis

AI agents—systems that operate autonomously to complete professional workflows—have emerged as a focus area for both researchers and enterprises seeking to automate complex, knowledge-based work. However, evaluating whether these agents actually perform reliably on real-world tasks has lacked a standard, industry-wide framework. Agents' Last Exam addresses this gap by building a comprehensive evaluation platform that measures agent performance against tasks drawn from actual professional practice across dozens of industries.

The involvement of 300+ industry experts suggests the benchmark is being anchored to work that practitioners recognize as meaningful and economically significant, rather than academic test cases disconnected from how agents will actually be deployed. By requiring verifiable outcomes and maintaining objective, comparable scoring across different professional domains, the benchmark aims to create a shared standard that both builders of AI agents and users evaluating them can rely on. The trajectory toward 5,000 tasks—moving from 1,500 currently collected—indicates this is a long-term effort to build breadth and depth across professional fields.

FAQ

What kinds of tasks are included in this benchmark?
The benchmark measures performance on long-horizon, economically valuable tasks with verifiable outcomes. It covers all 55 targeted sub-industries and spans most major fields of professional work performed on a computer.
How many tasks are in the benchmark now, and what is the goal?
The benchmark has collected 1,500+ tasks so far, with a target of 5,000 tasks total.

Get the latest Large Language Models news every morning

For example, today's edition would include:

  • Visko raises $10M, launches live AI video model OrbisSiliconANGLE AI · 1h ago
  • Runway unveils Solaris, an AI that generates app interfaces in real timeTHE DECODER · 1h ago
  • Google AI Search flags Facebook users as dangerTHE DECODER · 1h ago

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleEstonia's €24M Tax Blunder Sparked AI 'Fuckup Finder' for Draft Laws