AIToday
Large Language ModelsAI Business & IndustryLatent SpacePublished: Jul 25, 2026, 19:01 JST3 min read

Claude Opus 5 matches Fable performance at half the cost

Claude Opus 5 matches Fable performance at half the cost

3 Key Points

  1. What happened

    Anthropic released Claude Opus 5on Friday. On Artificial Analysis' agentic knowledge work benchmark (AA-Briefcase), Opus 5 outperformed Claude Fable 5 by nearly 150 Elo while reducing cost per task by 20%. Epoch reported Opus 5 achieved an ECI of 159—slightly below Fable 5's 161—but matched Fable 5 at 161 on software engineering benchmarks (SWE-ECI).

  2. Why it matters

    Independent benchmarks show Opus 5 delivers frontier-level capability at lower cost, challenging the traditional trade-off between performance and price. Users report strong practical improvements in coding and browser automation tasks, even where aggregate benchmark scores suggest only marginal gains. The launch underscores a market shift from static chat benchmarks toward agentic evaluations (tool use, browser control, software engineering loops) that traditional metrics may undervalue.

  3. What to watch

    Real-world leaderboard scores based on actual use are coming soon, according to Arena. Community evaluations are still catching up at posting time. Nous Research added Opus 5 access through its portal with a 20% discount applied to all models including Opus 5.

Ask the AI about this article →

Summaries like this, in your inbox every morning.

Context & Analysis

Claude Opus 5's launch reflects a fundamental shift in how frontier AI models are being evaluated and deployed. While Epoch's benchmark scores suggest only modest gains over Fable 5 at the aggregate level (159 ECI vs. 161), independent evaluations and user feedback indicate meaningful practical improvements—particularly in coding, software engineering tasks, and agentic workflows. The tension between benchmark scores and qualitative user reports points to a broader ecosystem problem: traditional capability indices compress diverse behaviors into single numbers, obscuring specialized strengths in domains like software engineering or tool use that users care most about.

The market context explains the launch's reception. Anthropic enters a crowded frontier field where cost efficiency now matters as much as raw capability, and where agentic competence (browser control, task automation, multi-step reasoning) has become a key competitive wedge. Users are no longer asking whether a model matches the previous frontier leader on a static chat benchmark; they are asking whether it can reliably execute real-world workflows—browser automation, code generation loops, and tool invocation—at acceptable cost. Opus 5's ability to deliver near-Fable performance at substantially lower cost addresses both dimensions of this shift.

Community response also reflects growing skepticism about whether published benchmarks keep pace with deployed capability. One user called the ECI result 'incredibly underrated,' arguing Opus 5 feels 'much better at everything' in practice despite scoring only 1 point higher than the prior Opus 4.8 model. This friction is not unique to Opus 5; it signals a wider challenge frontier labs face as test-time compute and search strategies complicate the relationship between static evals and real-world performance. Real-world leaderboards based on actual use are forthcoming, but until then, practitioner judgment and anecdote will likely outweigh aggregate benchmarks in shaping perception.

FAQ
How does Opus 5 perform on benchmarks compared to Fable 5?
Epoch reported Opus 5 achieved an ECI of 159, slightly below Fable 5's 161 overall, but matched Fable 5 at 161 on software engineering benchmarks (SWE-ECI). On Artificial Analysis' agentic knowledge work benchmark (AA-Briefcase), Opus 5 outperformed Fable 5 by nearly 150 Elo.
How much cheaper is Opus 5 than Fable?
On the AA-Briefcase benchmark, Opus 5 reduced cost per task by 20% compared to Fable 5. Nous Portal is also offering a 20% discount applied to all models including Opus 5.
What practical strengths are users reporting?
Users highlighted strong coding performance and agentic tool use, particularly browser automation and control. One user reported Opus 5 successfully opened a browser and canceled a ChatGPT Pro subscription, noting 'This thing can really drive a browser wow.'

Get the latest Large Language Models news every morning

For example, today's edition would include:

  • ChatGPT Work with GPT-6 Astra builds 5K running loop in 27 minutesSimon Willison's Weblog · 1h ago
  • OpenAI agents linked to RubyGems attack in MayThe Verge AI · 1h ago
  • Tesla Reportedly Pushes Staff Toward Grok 4.5 as AI Spending Cap Takes EffectTop Companies AI · 5h ago

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleClaude Opus 5 outperforms Fable 5, costs less on most benchmarks