
Anthropic's Claude Opus 5 has become the most capable AI model available according to several benchmarks, scoring 61 on the Artificial Analysis Intelligence Index and outperforming Fable 5 while costing less. On knowledge work tasks, Opus 5 costs $10.41 at its default "high" reasoning tier compared with $22.30 for Fable 5, while achieving superior analytical quality. The result underscores that frontier AI models are converging in capability—no single model can claim a clear advantage—lending weight to the argument that AI models will eventually become commoditized.
Summaries like this, in your inbox every morning.
Sign up free →What happened
Anthropic released Claude Opus 5, which scored 61 on the Artificial Analysis Intelligence Index—ahead of Fable 5 (60)—and matched or beat Fable 5 on most benchmarks including coding, scientific reasoning, and knowledge work tasks. Token pricing remains at $5 per million input tokens and $25 per million output tokens.
Why it matters
Opus 5 costs less than Fable 5 on average Intelligence Index tasks ($2.03 vs. $2.75) and substantially less on knowledge work—$10.41 at the "high" reasoning tier versus $22.30 for Fable 5—while delivering superior analytical quality (Analytical Quality Elo of 2016 vs. Fable 5's ~1700). This suggests frontier models are becoming more cost-competitive rather than pulling away from each other.
What to watch
The "high" reasoning tier emerges as the best value for Opus 5 on coding tasks, beating the costlier "max" tier due to time constraints limiting attempts. At "max," Opus 5 takes over 36 minutes per task on knowledge work, about 50% longer than Opus 4.8.
Anthropic's Claude Opus 5 has launched as the most capable AI model available today according to multiple independent benchmarks, though the gap between it and rivals like Fable 5 is now measured in single digits rather than substantial margins. On the Artificial Analysis Intelligence Index, which combines nine tests covering knowledge work, coding, scientific reasoning, and factual accuracy, Opus 5 scored 61 points, placing it just ahead of Fable 5 (60), GPT-5.6 Sol (59), Kimi K3 (57), and Claude Opus 4.8 (56). Artificial Analysis collaborated with Anthropic to test the model before its public release.
In specialized domains, Opus 5 demonstrates strengths and weaknesses. On the Artificial Analysis Coding Index, which measures how well AI models handle programming tasks independently, Opus 5 at "xhigh" paired with Claude Code shares first place on the leaderboard. On Terminal-Bench v2.1, a test of autonomous agents in real terminal environments, Opus 5 scored 89 percent at "max," matching the previous leader GPT-5.6 Sol. For scientific reasoning, Opus 5 scored 53 percent on Humanity's Last Exam and ties Fable 5 on CritPt, a physics benchmark from Argonne National Laboratory and UIUC researchers. However, factual accuracy remains a weakness: on AA-Omniscience, which tests the accuracy of a model's knowledge claims, Opus 5 improved by 7 points over Opus 4.8 but still trails Fable 5. Its hallucination rate stands at 50 percent, up 14 points from Opus 4.8. Epoch AI, a separate research institute, independently tested Opus 5 and assigned it an overall Epoch Capability Index score of 159, just below Fable 5 at 161; however, on software engineering benchmarks specifically, Opus 5 ties Fable 5 at 161.
Pricing and reasoning tiers emerge as Opus 5's key differentiation. Token pricing holds at $5 per million input tokens and $25 per million output tokens. On average Intelligence Index tasks, Opus 5 costs $2.03 compared with $2.75 for Fable 5. More significantly, on knowledge work—as measured by the AA-Briefcase benchmark, which tests typical office tasks like writing reports, building presentations, and analyzing spreadsheets from thousands of input files—Opus 5 at the "high" reasoning tier costs just $10.41 per task versus $22.30 for Fable 5, a drop of 53 percent. At "max" reasoning, Opus 5 costs $17.79, still about 20 percent below Fable 5. Opus 5's Elo rating on AA-Briefcase reaches 1720 at "max" reasoning, 146 points ahead of Fable 5 (1574). Its three highest reasoning tiers (max, xhigh, high) occupy the top three spots, and combined with Fable 5, Sonnet 5, and Opus 4.8, Anthropic models dominate the top-10 positions. On GDPval-AA v2, another knowledge-work test, Opus 5 at "max" reaches an Elo of 1861, far above the human baseline of 1000.
However, higher reasoning tiers do not always deliver proportional value. Research from Vals.ai tested Opus 5 across all five reasoning tiers using Vibe Code Bench, a programming benchmark. Scores climbed from 76.7 percent at "low" to 82 percent at "medium" and 89.8 percent at "high," but then dipped to 88.3 percent at "xhigh" and 88.4 percent at "max" despite much higher costs. Vals.ai found that the highest tiers tend to produce more complex solutions that contain errors more often, while the "high" tier produces simpler solutions meeting requirements more reliably. This pattern holds on Terminal-Bench 2.1 as well: the "high" tier beats "max" because the model spends more time on each attempt at the top tier, leaving fewer total attempts within the time limit. Anthropic has set "high" as the default tier in both the API and Claude Code, reflecting this practical insight. On knowledge work, Opus 5's biggest gains show up in analytical quality: at "max," it reaches an Analytical Quality Elo of 2016, nearly 300 points ahead of Fable 5. Its Rubric Pass Rate—how often it meets predefined quality standards—sits at 58 percent ("max"), 57.2 percent ("xhigh"), and 56 percent ("high"). Presentation quality is weaker: Opus 5 scores a Presentation Elo of 1628, about 40 points behind GPT-5.6 Sol at "max" (1666). Execution time on knowledge work is substantial: at "max," Opus 5 requires over 36 minutes per task and averages 103 passes, about 50 percent longer than Opus 4.8 at 24 minutes and 55 passes.
Claude Opus 5 arrives at a point of convergence among frontier AI models. Both Artificial Analysis and Epoch AI confirm that the performance gap between leading models has tightened considerably. Artificial Analysis tested Opus 5 at 61 points on its Intelligence Index, just one point ahead of Fable 5 and only two ahead of GPT-5.6 Sol. Epoch AI's independent assessment reinforces this picture: Opus 5 scores 159 on the overall Epoch Capability Index versus Fable 5 at 161, a margin too small to claim superiority. The article explicitly notes that "no single model can pull away or claim a clear advantage," suggesting that the era of incremental model releases each delivering decisive improvements may be ending.
What sets Opus 5 apart is not breakthrough capability but cost-effectiveness and specialization. On knowledge work—tasks like writing reports, building presentations, and analyzing data—Opus 5 achieves dramatic cost advantages. At the "high" reasoning tier it costs $10.41 per task versus $22.30 for Fable 5 while beating Fable 5 in Elo ranking. Its analytical quality is particularly strong, reaching an Analytical Quality Elo of 2016 at "max," nearly 300 points ahead of Fable 5. This pattern suggests that rather than all-purpose superiority, Opus 5 excels in specific domains and pricing configurations. The body's data on reasoning tiers reveals an interesting tension: higher reasoning tiers produce more complex (and error-prone) solutions, while the default "high" tier strikes a reliability-to-cost balance. This insight aligns with Anthropic's own API design, which sets "high" as the default—a practical acknowledgment that more computation does not always yield better results.
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytime
No comments yet. Be the first to share your thoughts!
Log in to join the discussion





Get curated AI news from 200+ sources delivered daily to your inbox. Free to use.
Get Started FreeFree · takes 30 seconds · unsubscribe anytime