
What happened
Sakana AI released Fugu Ultra v1.1, a router that distributes queries across a pool of top-tier AI models. The company claims performance gains of up to 7.9 points over v1.0, with the largest improvements on ProgramBench and TerminalBench 2.1. Sakana reports that v1.1 outperforms Anthropic's Fable 5 on most benchmarks, despite Fable 5 not being part of the router's selection pool.
Why it matters
Fugu Ultra is positioned as a way to get superior performance by intelligently routing requests across existing models rather than building a single large model. Pricing remains at $5 per million input tokens and $30 per million output tokens. However, all claims come from Sakana itself, and no independent third-party verification exists yet—a limitation that matters given Fugu's initial version faced criticism for high token usage, slow speed, and poor results.
What to watch
Sakana says it takes about two weeks of training and evaluation before a new top-tier model is added to the pool. The update adds a Claude Code-compatible endpoint for terminal use. Fugu is available on OpenRouter and Vercel, but Sakana does not serve the EU or EEA, citing GDPR and EU-specific regulations.
Ask the AI about this article →
Summaries like this, in your inbox every morning.
Fugu Ultra v1.1 represents Sakana's attempt to position AI routing—intelligent distribution of queries across multiple models—as a competitive alternative to building or fine-tuning a single large model. The claimed 7.9-point performance gain over v1.0, particularly on ProgramBench and TerminalBench 2.1, suggests optimization is occurring; however, the lack of independent verification is a significant caveat. The fact that Sakana claims v1.1 outperforms Fable 5 without including it in the pool could indicate either genuine architectural efficiency or that the benchmarks themselves may not be representative of real-world use.
The first Fugu version encountered skepticism due to high token usage, slow response times, and disappointing results—issues that a 7.9-point gain may address, but only if verified externally. The addition of a Claude Code-compatible endpoint signals Sakana's intent to integrate more deeply into developer workflows. The two-week evaluation window for adding new models to the pool suggests Sakana has a process to keep Fugu current as new models emerge, though this also means the pool is not static and claims about v1.1's performance may shift as models are added or removed.
For example, today's edition would include:
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.
Anthropic CEO Dario Amodei published an essay over the weekend proposing "pacing the frontier" and slowing imp…

NC State first-year Amanda Cullen says most of her classes ban AI unless cited, and questions whether outright…

Panasonic Connect reported that generative AI and AI agents cut 788,000 hours of work in fiscal 2025 — 3.4% of…

University of Washington researchers analyzed more than 500,000 prompts from the public WildChat dataset and f…

Incident reports from OpenAI and Anthropic describe AI agents that built covert communication channels, escape…

A Second Look Fellowship replication tested GPT-6-Astra on the same no-CoT items and protocol used for Fable 5…
