
Anthropic's Claude Opus 5 has achieved a major leap on ARC-AGI-3, a benchmark designed to test general reasoning on novel tasks, scoring 30.2 percent compared to the previous record of 7.8 percent by OpenAI's GPT-5.6 Sol (Max). ARC Prize attributes the gain to stronger logical reasoning that enables exploration and planning in unfamiliar environments, and the model demonstrated unprecedented behavior such as translating tasks into algebraic notation. However, performance on independent benchmarks like Witness suggests the gains may be narrower in real-world scenarios, indicating the improvement could reflect training tailored to ARC-AGI-3's specific puzzle formats.
Summaries like this, in your inbox every morning.
Sign up free →What happened
Anthropic's Claude Opus 5 scored 30.2 percent on ARC-AGI-3, a benchmark that measures how well AI models solve new tasks they didn't encounter during training. The previous record was 7.8 percent, set by OpenAI's GPT-5.6 Sol (Max). Opus 5 solved five previously unsolved environments, four of them at or above human level, and outperformed Anthropic's own Fable-class models, which hit around 20 percent.
Why it matters
ARC Prize creators attribute the lead to stronger logical reasoning that enables autonomous exploration, planning, and execution across unfamiliar environments. Opus 5 exhibited reasoning behavior researchers hadn't seen before—translating tasks into algebraic notation and independently formulating reflection equations for the first time. However, independent tests on Witness, a private benchmark for interactive puzzle games, suggest narrower real-world gains; Opus 5 scored 43.4 there, statistically tying competitors and improving far less over Opus 4.8 than it did on ARC-AGI-3.
What to watch
Six of the 25 public demo environments on ARC-AGI-3 have now been solved. Full results, replays, and benchmarking code are publicly available. On the older ARC-AGI-2 benchmark, Opus 5 scores 90.4 percent, and it reaches 97.5 percent on ARC-AGI-1. Researchers note that Opus 5 would likely score even higher if used within Claude Code, an external harness, though official scores count only the language model's own performance.
Anthropic's Claude Opus 5 has set a new high score on ARC-AGI-3, achieving 30.2 percent accuracy—a massive jump from the previous record of 7.8 percent set by OpenAI's GPT-5.6 Sol (Max). In doing so, Opus 5 solved five environments that had not been solved before, with four of them reaching or exceeding human-level performance. The model also surpassed Anthropic's own Fable-class models, which scored around 20 percent on the same benchmark. ARC Prize, the organization behind the benchmark, published full results, replays, and benchmarking code publicly, with six of the 25 public demo environments now solved.
ARC-AGI-3 is designed to measure how well AI models can handle genuinely novel tasks—ones they never saw during training, including tasks that humans typically solve with ease. Unlike benchmarks that reward stored knowledge, ARC-AGI-3 works as an interactive game: the model must infer the rules of an environment, plan a sequence of actions, and execute them step by step. This format tests general reasoning ability and the capacity to adapt to unfamiliar patterns. According to ARC Prize's analysis, Opus 5's commanding lead stems from stronger logical reasoning that "enables more autonomous exploration, planning, and execution across unfamiliar environments." The model demonstrated unprecedented behavior during testing—translating tasks into algebraic notation and independently formulating reflection equations for the first time, a capability researchers had not observed in prior models.
However, the source of Opus 5's improvement remains unclear and contested by independent researchers. Anthropic has not disclosed how the model was trained or optimized. Targeted data labeling and reinforcement learning are plausible mechanisms: annotators could have labeled reasoning traces, useful actions, failed attempts, and recovery steps from similar puzzles, with reinforcement learning then rewarding exploration, planning, rule discovery, and self-correction. Notably, Opus 5 was developed after ARC-AGI-3 and its format became public, which may have allowed Anthropic to tailor training to the benchmark's specific skills and puzzle types—though the company has not shown it trained on the exact tasks themselves.
Independent testing suggests a more cautious picture. Guanghan Ning, who designed Witness—a private benchmark for interactive puzzle games—found that Opus 5 scored 43.4, statistically tying competitors like Kimi K3 and Fable 5 while improving far less over Opus 4.8 than it did on ARC-AGI-3. In one test, Opus 5 identified a conventional puzzle's hidden rules before taking any action but trailed Opus 4.8 on a game with less familiar mechanics. Ning suggests the pattern fits training on genre-specific data, though Witness cannot identify what data Anthropic used. Greg Kamradt, one of the ARC-AGI-3 researchers, countered that a single weak result does not outweigh the model's overall improvement and that Witness—designed around ARC-AGI-3-style puzzles—may not fully test the kind of novelty adaptation that ARC-AGI-3 measures. Ning later clarified that Opus 5 did generalize to Witness, just far less than on ARC-AGI-3, and drew a parallel to how coding benchmarks evolved: as a major target for interactive reasoning, ARC-AGI-3 will likely attract the most training effort first, and covering more edge cases could help models generalize to a wider range of abstract reasoning tasks.
On older versions of the benchmark, Opus 5's performance held steady: it scores 90.4 percent on ARC-AGI-2 and 97.5 percent on ARC-AGI-1, matching previous top scores though at slightly higher costs. ARC Prize also notes that some AI systems may have already passed the benchmark when used with external software tools called a harness, but official scores count only the language model's own performance. Opus 5 would likely score even higher if used within Claude Code, such a harness, but the official leaderboard reflects the model operating independently.
Claude Opus 5's performance on ARC-AGI-3 represents a dramatic leap—jumping from 7.8 percent to 30.2 percent—but the nature of that improvement remains contested. ARC Prize credits stronger logical reasoning capabilities that enable more autonomous exploration and planning in unfamiliar environments. The model's ability to translate tasks into algebraic notation and independently formulate reflection equations marks genuinely novel behavior from a reasoning standpoint. However, Anthropic has not disclosed its training approach, leaving open the question of whether the gain reflects fundamental reasoning improvements or targeted optimization for ARC-AGI-3's puzzle-specific mechanics.
The tension between these interpretations comes into focus when comparing results across benchmarks. Independent testing on Witness, another interactive puzzle benchmark, shows Opus 5 scoring 43.4—statistically equivalent to competitors like Kimi K3 and Fable 5, with marginal gains over Opus 4.8 compared to its massive ARC-AGI-3 jump. Guanghan Ning, who designed Witness, suggests this pattern fits training on genre-specific data, though he clarifies that Opus 5 did generalize to Witness, just far less dramatically. Greg Kamradt, a researcher behind ARC-AGI-3, counters that Witness may not fully test novelty adaptation and that broader reasoning gains are not ruled out by a single weak result.
The resolution may lie in how AI benchmarking evolves. Ning draws a parallel to coding benchmarks, which moved from saturated benchmarks like HumanEval to frequently updated competitions as models improved. With ARC-AGI-3 now emerging as a major target for interactive reasoning research, it will likely attract concentrated training effort first; covering more edge cases and generalization patterns could then help models transfer those skills to wider reasoning tasks. Until independent benchmarks that test true novelty-handling mature, the question of whether Opus 5's ARC-AGI-3 lead reflects fundamental progress or focused optimization will remain open.
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytime
No comments yet. Be the first to share your thoughts!
Log in to join the discussion





Get curated AI news from 200+ sources delivered daily to your inbox. Free to use.
Get Started FreeFree · takes 30 seconds · unsubscribe anytime