
What happened
Anthropic's Claude Opus 5 scored 30.2 percent on ARC-AGI-3, a benchmark that measures how well AI models solve new tasks they didn't encounter during training. The previous record was 7.8 percent, set by OpenAI's GPT-5.6 Sol (Max). Opus 5 solved five previously unsolved environments, four of them at or above human level, and outperformed Anthropic's own Fable-class models, which hit around 20 percent.
Why it matters
ARC Prize creators attribute the lead to stronger logical reasoning that enables autonomous exploration, planning, and execution across unfamiliar environments. Opus 5 exhibited reasoning behavior researchers hadn't seen before—translating tasks into algebraic notation and independently formulating reflection equations for the first time. However, independent tests on Witness, a private benchmark for interactive puzzle games, suggest narrower real-world gains; Opus 5 scored 43.4 there, statistically tying competitors and improving far less over Opus 4.8 than it did on ARC-AGI-3.
What to watch
Six of the 25 public demo environments on ARC-AGI-3 have now been solved. Full results, replays, and benchmarking code are publicly available. On the older ARC-AGI-2 benchmark, Opus 5 scores 90.4 percent, and it reaches 97.5 percent on ARC-AGI-1. Researchers note that Opus 5 would likely score even higher if used within Claude Code, an external harness, though official scores count only the language model's own performance.
Summaries like this, in your inbox every morning.
Claude Opus 5's performance on ARC-AGI-3 represents a dramatic leap—jumping from 7.8 percent to 30.2 percent—but the nature of that improvement remains contested. ARC Prize credits stronger logical reasoning capabilities that enable more autonomous exploration and planning in unfamiliar environments. The model's ability to translate tasks into algebraic notation and independently formulate reflection equations marks genuinely novel behavior from a reasoning standpoint. However, Anthropic has not disclosed its training approach, leaving open the question of whether the gain reflects fundamental reasoning improvements or targeted optimization for ARC-AGI-3's puzzle-specific mechanics.
The tension between these interpretations comes into focus when comparing results across benchmarks. Independent testing on Witness, another interactive puzzle benchmark, shows Opus 5 scoring 43.4—statistically equivalent to competitors like Kimi K3 and Fable 5, with marginal gains over Opus 4.8 compared to its massive ARC-AGI-3 jump. Guanghan Ning, who designed Witness, suggests this pattern fits training on genre-specific data, though he clarifies that Opus 5 did generalize to Witness, just far less dramatically. Greg Kamradt, a researcher behind ARC-AGI-3, counters that Witness may not fully test novelty adaptation and that broader reasoning gains are not ruled out by a single weak result.
The resolution may lie in how AI benchmarking evolves. Ning draws a parallel to coding benchmarks, which moved from saturated benchmarks like HumanEval to frequently updated competitions as models improved. With ARC-AGI-3 now emerging as a major target for interactive reasoning research, it will likely attract concentrated training effort first; covering more edge cases and generalization patterns could then help models transfer those skills to wider reasoning tasks. Until independent benchmarks that test true novelty-handling mature, the question of whether Opus 5's ARC-AGI-3 lead reflects fundamental progress or focused optimization will remain open.
For example, today's edition would include:
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.
SoftBank Group is seeking the equivalent of more than $11 billion, issuing $10 billion of dollar securities ac…

Google DeepMind, part of Alphabet, has been named in a new class-action lawsuit challenging an industry-funded…

In the NEXER Group and LISKILLING survey, 90.7% said they had no workplace AI or DX training experience, while…

Google confirmed that Gemini connected to the internet and accessed the systems of three real companies during…

Nvidia guided to $108 billion in quarterly revenue, up from $96.2 billion

Google will invest at least €13 billion ($15.1 billion) in Finland over the next two years for three new data…
