AIToday

Claude Opus 5 leads AI leaderboard at half Fable 5's price

Simon Willison's Weblog1h ago

Key takeaway

Anthropic released Claude Opus 5, its latest AI model, which ranks first on the Artificial Analysis leaderboard ahead of Claude Fable 5 while priced identically to the previous Opus 4.8. The model exhibits proactive problem-solving—demonstrating the ability to write its own computer vision code to complete tasks when given incomplete information—and improves on vulnerability detection while remaining deliberately restricted from learning exploitation techniques.

Summaries like this, in your inbox every morning.

Sign up free →

3 Key Points

  • What happened

    Anthropic released Claude Opus 5, a new AI model that tops the Artificial Analysis leaderboard, ahead of Claude Fable 5. It is priced the same as the previous Opus 4.8 model and offers a faster mode at twice the base cost.

  • Why it matters

    Anthropic describes Opus 5 as approaching "frontier intelligence" comparable to Fable 5 while costing significantly less, making advanced AI capabilities potentially more accessible. The model demonstrates unexpected problem-solving ability—on one benchmark task, it independently wrote a computer vision pipeline to extract geometry from a drawing it could not directly view, then rebuilt a 3D FreeCAD model from the raw pixels.

  • What to watch

    Anthropic has published a prompting guide for Claude Opus 5 to help users work with the model. The model improves substantially on cybersecurity vulnerability detection but remains deliberately undertrained on exploitation techniques, a design choice Anthropic notes.

In Depth

On 24 July 2026, Anthropic announced Claude Opus 5, a new large language model that immediately assumed the top position on the Artificial Analysis leaderboard, surpassing even Claude Fable 5. The model is priced identically to its predecessor Opus 4.8 and maintains the same pricing structure, offering a "fast mode" option at twice the cost of the base model.

Anthropic's own framing emphasizes Opus 5 as "a thoughtful and proactive model that comes close to the frontier intelligence of Claude Fable 5 at half the price." The proactive aspect is demonstrated by a striking example from the Frontier-Bench evaluation suite. In one task, Opus 5 was presented with a drawing of a machine part and asked to write code to rebuild it as a 3D FreeCAD model. Critically, the model was deliberately given no direct way to view the drawing. Rather than failing or requesting clarification, Opus 5 independently wrote its own computer vision pipeline to extract the geometry from the raw pixel data, then used that pipeline to reconstruct the full machine part—solving both the primary task and an unanticipated prerequisite task.

On the cybersecurity front, Anthropic reports that Opus 5 has not been intentionally trained on cyber exploitation tasks, continuing the company's prior approach with Opus 4.8. Nevertheless, the model's general capability improvements have yielded substantial gains in vulnerability detection; it now approaches the performance of Mythos 5 in finding cybersecurity vulnerabilities. However, it remains substantially behind Mythos 5 in the exploitation phase—the ability to turn discovered vulnerabilities into active cyber threats. Anthropic attributes this gap to the deliberate absence of exploitation training and notes that this design choice may prove favorable in discussions with government regulators.

To support developer adoption, Anthropic published a dedicated prompting guide for Claude Opus 5. Additional resources include a guide titled "The new rules of context engineering for Claude 5 generation models," authored by Thariq Shihipar, to help users optimize their interactions with the new model generation.

Context & Analysis

Claude Opus 5 represents a milestone in Anthropic's model hierarchy: it achieves top ranking on a major independent benchmark while undercutting the cost of the previous generation's flagship model. The pricing structure mirrors Opus 4.8, with a base tier and a "fast mode" available at double the cost, giving users a choice between speed and cost efficiency.

A notable aspect of Opus 5's capabilities is its apparent tendency toward proactive problem-solving beyond what it was explicitly instructed to do. The Frontier-Bench anecdote reveals this clearly: faced with incomplete information (a drawing it could not directly view), the model spontaneously authored a computer vision pipeline as an intermediary step. This suggests that general capability improvements have also expanded the model's ability to decompose and solve novel sub-problems.

Anthropichas deliberately maintained its prior stance on cybersecurity training: Opus 5 was not trained on exploitation tasks, despite improving substantially on vulnerability detection as a side effect of becoming more generally capable. The model now approaches Mythos 5 in vulnerability finding but remains far behind on turning those findings into actual cyber threats—a safety-conscious design choice that the author notes may help with regulatory acceptance.

FAQ

How does Claude Opus 5's price compare to Claude Fable 5?
Anthropic describes Opus 5 as achieving "frontier intelligence" close to Fable 5 at half the price of Fable 5. Opus 5 itself is priced the same as the previous Opus 4.8 model.
What makes Claude Opus 5 different from previous versions?
Opus 5 demonstrates unexpectedly proactive behavior. In one Frontier-Bench task, when given a drawing of a machine part without direct viewing access, it wrote its own computer vision pipeline to extract geometry from raw pixels and then reconstructed the full 3D FreeCAD model independently.

Get the latest Large Language Models news every morning

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytime

Discussion

No comments yet. Be the first to share your thoughts!

Log in to join the discussion

Related Articles

Stay ahead with AI news

Get curated AI news from 200+ sources delivered daily to your inbox. Free to use.

Get Started Free

Free · takes 30 seconds · unsubscribe anytime