AIToday
Large Language ModelsAI Business & IndustrySiliconANGLE AIPublished: Sep 11, 2026, 10:01 JST2 min read

DeepSeek-V4.1-Flash outruns V4-Pro, reroutes its API calls

DeepSeek-V4.1-Flash outruns V4-Pro, reroutes its API calls

3 Key Points

  1. What happened

    DeepSeek released DeepSeek-V4.1-Flash, the smallest model in a new architecture family, and says outside tests put it ahead of its larger flagship V4-Pro on performance, cost, speed and total runtime.

  2. Why it matters

    From Sept. 14, V4-Pro API requests are answered by V4.1-Flash at the smaller model's rates until a V4.1-Pro launches, cutting peak output pricing from $3.96 to $1.20 per million tokens, roughly 70%.

  3. What to watch

    Whether developers accept being moved off V4-Pro, since V4-Flash and August's experimental vision model are now retired and their calls already land on V4.1-Flash. Watch the V4.1-Pro launch date.

WHO IT HITSDevelopers and API teams that build on DeepSeek models see their V4-Pro output costs fall roughly 70% from Sept. 14, while teams relying on the retired V4-Flash or the August experimental vision model are shifted to V4.1-Flash without a choice.

Ask the AI about this article →

Summaries like this, in your inbox every morning.

Context & Analysis

DeepSeek's pitch here is unusual: the smallest member of a new architecture family is being positioned above the company's own flagship. V4.1-Flash is a mixture-of-experts model with 552 billion parameters, close to double the 284 billion in V4-Flash, but a new causal encoder-decoder design keeps only 8 billion parameters active while reading a prompt and 16 billion while generating output. Much of the engineering went into shrinking the key-value cache, which the technical report says is stored in a four-bit floating-point format at 890 bytes per token, about a quarter of what V4-Flash needs, with persistent SSD cache storage down to roughly an eighth of the previous generation.

The pricing move is the practical part. Rather than ask developers to switch, DeepSeek is switching them: from Sept. 14, V4-Pro API requests are answered by V4.1-Flash and billed at the smaller model's rates until a V4.1-Pro version launches, and the retired V4-Flash and August's experimental vision model already route there. Image understanding, previously only in that experimental release, is now built into the model itself.

The launch lands the same day Anthropic named DeepSeek in its threat intelligence report as one of seven China-based labs it says ran distillation campaigns against Claude, attributing more than 12.1 million exchanges over 14 days in July to DeepSeek. DeepSeek's ability to hold developers on the rerouted API may hinge on whether the promised V4.1-Pro arrives soon, and on whether buyers weigh Anthropic's claims alongside the roughly 70% output-price cut.

FAQ
How much does DeepSeek-V4.1-Flash cost?
Off-peak API pricing is 15 cents per million uncached input tokens and 60 cents per million output tokens, with rates doubling during weekday peak windows.
What happens to my V4-Pro requests after Sept. 14?
They will be answered by V4.1-Flash and billed at the smaller model's rates until a V4.1-Pro version launches.
Where can I get the model weights?
The weights are available on Hugging Face under the MIT license, and the model is live in DeepSeek's web and mobile apps.
How does it compare with Claude Opus 5 and GPT-5.6 Sol?
DeepSeek's benchmark table shows V4.1-Flash scored 90.6 on Terminal-Bench 2.1, ahead of Opus 5 at 89.1 and GPT-5.6 Sol at 88.8, though both U.S. models still lead on GPQA Diamond.
SiliconANGLE AIRead Original Article

Get the latest Large Language Models news every morning

For example, today's edition would include:

  • Dynatrace acquires Arize AI as observability shifts to actionSiliconANGLE AI · 4h ago
  • Shared base cuts 100 fine-tunes from 1.5 TB to 19.3 GBDaily Dose of Data Science · 4h ago
  • OpenAI agents hit RubyGems, undisclosed since May 12thSimon Willison's Weblog · 4h ago

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleGlobal IC substrate market to surge 32% to US$19.5 billion in 2026