AIToday

Sakana's Fugu Ultra v1.1 claims to beat Claude's Fable 5 without using it

THE DECODER3h ago
Sakana's Fugu Ultra v1.1 claims to beat Claude's Fable 5 without using it

Key takeaway

Sakana AI released Fugu Ultra v1.1, an AI router that distributes queries across a pool of public models and claims to outperform Anthropic's Claude Fable 5, even though Fable 5 is not part of its model pool. The update shows gains of up to 7.9 points over the previous version on benchmarks like ProgramBench and TerminalBench 2.1. Pricing stays at $5 per million input tokens and $30 per million output tokens. However, all performance claims come from Sakana and have not been independently verified.

Summaries like this, in your inbox every morning.

Sign up free →

3 Key Points

  • What happened

    Sakana AI released Fugu Ultra v1.1, a router that distributes queries across a pool of top-tier AI models. The company claims performance gains of up to 7.9 points over v1.0, with the largest improvements on ProgramBench and TerminalBench 2.1. Sakana reports that v1.1 outperforms Anthropic's Fable 5 on most benchmarks, despite Fable 5 not being part of the router's selection pool.

  • Why it matters

    Fugu Ultra is positioned as a way to get superior performance by intelligently routing requests across existing models rather than building a single large model. Pricing remains at $5 per million input tokens and $30 per million output tokens. However, all claims come from Sakana itself, and no independent third-party verification exists yet—a limitation that matters given Fugu's initial version faced criticism for high token usage, slow speed, and poor results.

  • What to watch

    Sakana says it takes about two weeks of training and evaluation before a new top-tier model is added to the pool. The update adds a Claude Code-compatible endpoint for terminal use. Fugu is available on OpenRouter and Vercel, but Sakana does not serve the EU or EEA, citing GDPR and EU-specific regulations.

In Depth

Sakana AI has released Fugu Ultra v1.1, an updated version of its AI router—a system that distributes incoming queries across a curated pool of publicly available top-tier models rather than relying on a single monolithic model. The company claims the update delivers performance gains of up to 7.9 points relative to v1.0, with the most significant improvements appearing on ProgramBench and TerminalBench 2.1. In a claim that could attract attention, Sakana reports that Fugu Ultra v1.1 outperforms Anthropic's Fable 5 on most benchmarks, even though Fable 5 is not included in the router's model pool.

It is important to note that these performance numbers come entirely from Sakana and have not yet been verified by independent third parties. The pricing model for Fugu Ultra v1.1 remains unchanged from prior versions: $5 per million input tokens and $30 per million output tokens. Sakana indicates that when a new top-tier model becomes available, it undergoes about two weeks of training and evaluation before being integrated into Fugu's selection pool. The technical details of this update are documented in Sakana's technical report. In addition to performance improvements, v1.1 introduces a Claude Code-compatible endpoint, allowing developers to call Fugu directly from the terminal.

Fugu has been accessible via platforms such as OpenRouter and Vercel since its launch. The earlier version of Fugu received mixed to negative feedback from users and critics, who highlighted concerns about high token consumption, slow response speeds, and weak results. A regional limitation applies: Sakana does not currently serve the EU or EEA, a restriction the company attributes to GDPR and other EU-specific regulatory requirements.

Context & Analysis

Fugu Ultra v1.1 represents Sakana's attempt to position AI routing—intelligent distribution of queries across multiple models—as a competitive alternative to building or fine-tuning a single large model. The claimed 7.9-point performance gain over v1.0, particularly on ProgramBench and TerminalBench 2.1, suggests optimization is occurring; however, the lack of independent verification is a significant caveat. The fact that Sakana claims v1.1 outperforms Fable 5 without including it in the pool could indicate either genuine architectural efficiency or that the benchmarks themselves may not be representative of real-world use.

The first Fugu version encountered skepticism due to high token usage, slow response times, and disappointing results—issues that a 7.9-point gain may address, but only if verified externally. The addition of a Claude Code-compatible endpoint signals Sakana's intent to integrate more deeply into developer workflows. The two-week evaluation window for adding new models to the pool suggests Sakana has a process to keep Fugu current as new models emerge, though this also means the pool is not static and claims about v1.1's performance may shift as models are added or removed.

FAQ

How much does Fugu Ultra v1.1 cost?
Pricing is $5 per million input tokens and $30 per million output tokens.
How long does it take to add a new model to Fugu's pool?
Sakana says it takes about two weeks of training and evaluation before a new top-tier model is added to the pool.
Where is Fugu available?
Fugu is available on platforms like OpenRouter and Vercel. Sakana does not serve the EU or EEA, citing GDPR and EU-specific regulations.

Get the latest Large Language Models news every morning

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytime

Discussion

No comments yet. Be the first to share your thoughts!

Log in to join the discussion

Related Articles

Stay ahead with AI news

Get curated AI news from 200+ sources delivered daily to your inbox. Free to use.

Get Started Free

Free · takes 30 seconds · unsubscribe anytime