AIToday
Large Language ModelsAI Business & IndustryTHE DECODERPublished: Jul 26, 2026, 22:00 JST

Anthropic's Opus 5 reaches 30.2% on ARC-AGI-3, crushing prior 7.8% record

Anthropic's Opus 5 reaches 30.2% on ARC-AGI-3, crushing prior 7.8% record

3 Key Points

  1. What happened

    Anthropic's Claude Opus 5 scored 30.2 percent on ARC-AGI-3, a benchmark that measures how well AI models solve new tasks they didn't encounter during training. The previous record was 7.8 percent, set by OpenAI's GPT-5.6 Sol (Max). Opus 5 solved five previously unsolved environments, four of them at or above human level, and outperformed Anthropic's own Fable-class models, which hit around 20 percent.

  2. Why it matters

    ARC Prize creators attribute the lead to stronger logical reasoning that enables autonomous exploration, planning, and execution across unfamiliar environments. Opus 5 exhibited reasoning behavior researchers hadn't seen before—translating tasks into algebraic notation and independently formulating reflection equations for the first time. However, independent tests on Witness, a private benchmark for interactive puzzle games, suggest narrower real-world gains; Opus 5 scored 43.4 there, statistically tying competitors and improving far less over Opus 4.8 than it did on ARC-AGI-3.

  3. What to watch

    Six of the 25 public demo environments on ARC-AGI-3 have now been solved. Full results, replays, and benchmarking code are publicly available. On the older ARC-AGI-2 benchmark, Opus 5 scores 90.4 percent, and it reaches 97.5 percent on ARC-AGI-1. Researchers note that Opus 5 would likely score even higher if used within Claude Code, an external harness, though official scores count only the language model's own performance.

Not sure about something? Ask the AI

Summaries like this, in your inbox every morning.

Context & Analysis

Claude Opus 5's performance on ARC-AGI-3 represents a dramatic leap—jumping from 7.8 percent to 30.2 percent—but the nature of that improvement remains contested. ARC Prize credits stronger logical reasoning capabilities that enable more autonomous exploration and planning in unfamiliar environments. The model's ability to translate tasks into algebraic notation and independently formulate reflection equations marks genuinely novel behavior from a reasoning standpoint. However, Anthropic has not disclosed its training approach, leaving open the question of whether the gain reflects fundamental reasoning improvements or targeted optimization for ARC-AGI-3's puzzle-specific mechanics.

The tension between these interpretations comes into focus when comparing results across benchmarks. Independent testing on Witness, another interactive puzzle benchmark, shows Opus 5 scoring 43.4—statistically equivalent to competitors like Kimi K3 and Fable 5, with marginal gains over Opus 4.8 compared to its massive ARC-AGI-3 jump. Guanghan Ning, who designed Witness, suggests this pattern fits training on genre-specific data, though he clarifies that Opus 5 did generalize to Witness, just far less dramatically. Greg Kamradt, a researcher behind ARC-AGI-3, counters that Witness may not fully test novelty adaptation and that broader reasoning gains are not ruled out by a single weak result.

The resolution may lie in how AI benchmarking evolves. Ning draws a parallel to coding benchmarks, which moved from saturated benchmarks like HumanEval to frequently updated competitions as models improved. With ARC-AGI-3 now emerging as a major target for interactive reasoning research, it will likely attract concentrated training effort first; covering more edge cases and generalization patterns could then help models transfer those skills to wider reasoning tasks. Until independent benchmarks that test true novelty-handling mature, the question of whether Opus 5's ARC-AGI-3 lead reflects fundamental progress or focused optimization will remain open.

FAQ
What is ARC-AGI-3 and why does it matter?
ARC-AGI-3 measures how well AI models solve new tasks they didn't encounter during training, including ones humans can usually handle with ease. The current version works like a game in which the model must infer the rules of an interactive environment, plan its actions, and carry them out step by step, testing general reasoning rather than stored knowledge.
How much better did Opus 5 perform than previous records?
Opus 5 scored 30.2 percent on ARC-AGI-3, compared to the previous record of 7.8 percent set by GPT-5.6 Sol (Max). It also solved five previously unsolved environments, four of them at or above human level.
Do independent tests confirm the same level of improvement?
No. On Witness, a private benchmark for interactive puzzle games, Opus 5 scored 43.4, statistically tying Kimi K3 and Fable 5 while improving far less over Opus 4.8 than it did on ARC-AGI-3, suggesting the gains may be more modest on novel reasoning tasks outside ARC-AGI-3's specific formats.

Get the latest Large Language Models news every morning

For example, today's edition would include:

  • Google's Gemini broke into 3 real firms in a test, WSJ reportsITmedia AI+ · 6h ago
  • Nvidia's Huang: AI in "high production ramp" as revenue hits $108 billionYahoo Finance AI · 6h ago
  • llm-keys-ui 0.1 keeps API keys out of agent chatsSimon Willison's Weblog · 6h ago

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleTalent debt from AI is now visible in workforce data