AIToday
Large Language ModelsImage GenerationAI Business & IndustrySimon Willison's WeblogPublished: Sep 2, 2026, 10:00 JST2 min read

Anthropic's Fable 5.1: Costlier reasoning, better pelicans

Anthropic's Fable 5.1: Costlier reasoning, better pelicans

Key takeaway

  • Anthropic released Claude Fable 5.1, which scores highly on a new science benchmark.

  • In a creative test, its most expensive reasoning mode produced the best pelican image yet.

  • This mode cost $3.30 and took almost 14 minutes.

3 Key Points

  1. What happened

    Anthropic released Claude Fable (and Mythos) 5.1 on September 1st, 2026. In a test, the model's 'max' reasoning setting produced the best AI-generated pelican image the author has seen from Anthropic, at a cost of $3.30 and 13 minutes 54 seconds of processing.

  2. Why it matters

    The model scored 52.6% on the new Terminal-Bench-Science 0.1 benchmark, up from 24.7% for Fable 5, 29.0% for Opus 5, and 22.4% for GPT-5.6 Sol. However, the author notes that at lower reasoning settings ('low' and 'medium'), the model appeared to skip reasoning entirely for the pelican prompt, costing around 10 cents.

  3. What to watch

    The 'max' reasoning level produced 65,927 output tokens and a detailed reasoning trace, leading to a pelican with a blue hat and a basket with a fish. When this output was animated with another prompt, it cost $1.37 and produced a video where the wheels rotated correctly in the original SVG.

Ask the AI about this article →

Context & Analysis

The release of Claude Fable 5.1 is highlighted by a significant jump on a brand-new scientific benchmark, Terminal-Bench-Science 0.1, where it scored 52.6% compared to 24.7% for its predecessor. This suggests a notable improvement in handling scientific research tasks, a key area for Anthropic's marketing.

However, the author's hands-on testing reveals a more nuanced story about the model's reasoning capabilities. The cost and time for generating a simple SVG varied dramatically with the reasoning effort setting, from about 10 cents and 24 seconds at 'low' to $3.30 and nearly 14 minutes at 'max'. Intriguingly, the lower efforts seemed to bypass reasoning entirely for this task.

The 'max' setting's lengthy reasoning trace shows the model engaging in detailed, deliberate planning, from checking visual collisions to fixing a bicycle fork's curve. This suggests that for complex, iterative tasks, the higher cost tiers provide a genuinely different and more capable mode of operation, though the author notes it still lacks the creative flair of a competing model like Gemini 3.7 Flash.

FAQ

How much does the most detailed reasoning mode cost?
The 'max' reasoning effort for one SVG generation prompt cost $3.30 and took 13 minutes and 54 seconds. It used 65,927 output tokens.
What is special about the new Terminal-Bench-Science benchmark score?
Fable 5.1 scored 52.6% on the brand new Terminal-Bench-Science 0.1 benchmark. This is up from 24.7% for Fable 5, 29.0% for Opus 5, and 22.4% for GPT-5.6 Sol.
Did the model perform well at all reasoning levels?
No. For the pelican prompt, the 'low' and 'medium' settings appeared to skip reasoning entirely, producing similar results. Significant improvement was only seen at the 'xhigh' and 'max' settings.
Simon Willison's WeblogRead Original Article

Get the latest Large Language Models news every morning

For example, today's edition would include:

  • Nvidia invests $3.5B in MediaTek to profit from custom AI chipsYahoo Finance AI · 2h ago
  • Paid actors, AI scripts: Viral anti-Democrat YouTube network exposedSemafor Tech · 2h ago
  • ICRA panel warns of paper floodRobohub · 2h ago

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleAnthropic's Claude Fable 5.1 Now on Snowflake Cortex AI