
What happened
DeepSeek released DeepSeek-V4.1-Flash, the smallest model in a new architecture family, and says outside tests put it ahead of its larger flagship V4-Pro on performance, cost, speed and total runtime.
Why it matters
From Sept. 14, V4-Pro API requests are answered by V4.1-Flash at the smaller model's rates until a V4.1-Pro launches, cutting peak output pricing from $3.96 to $1.20 per million tokens, roughly 70%.
What to watch
Whether developers accept being moved off V4-Pro, since V4-Flash and August's experimental vision model are now retired and their calls already land on V4.1-Flash. Watch the V4.1-Pro launch date.
WHO IT HITSDevelopers and API teams that build on DeepSeek models see their V4-Pro output costs fall roughly 70% from Sept. 14, while teams relying on the retired V4-Flash or the August experimental vision model are shifted to V4.1-Flash without a choice.
Ask the AI about this article →
Summaries like this, in your inbox every morning.
DeepSeek's pitch here is unusual: the smallest member of a new architecture family is being positioned above the company's own flagship. V4.1-Flash is a mixture-of-experts model with 552 billion parameters, close to double the 284 billion in V4-Flash, but a new causal encoder-decoder design keeps only 8 billion parameters active while reading a prompt and 16 billion while generating output. Much of the engineering went into shrinking the key-value cache, which the technical report says is stored in a four-bit floating-point format at 890 bytes per token, about a quarter of what V4-Flash needs, with persistent SSD cache storage down to roughly an eighth of the previous generation.
The pricing move is the practical part. Rather than ask developers to switch, DeepSeek is switching them: from Sept. 14, V4-Pro API requests are answered by V4.1-Flash and billed at the smaller model's rates until a V4.1-Pro version launches, and the retired V4-Flash and August's experimental vision model already route there. Image understanding, previously only in that experimental release, is now built into the model itself.
The launch lands the same day Anthropic named DeepSeek in its threat intelligence report as one of seven China-based labs it says ran distillation campaigns against Claude, attributing more than 12.1 million exchanges over 14 days in July to DeepSeek. DeepSeek's ability to hold developers on the rerouted API may hinge on whether the promised V4.1-Pro arrives soon, and on whether buyers weigh Anthropic's claims alongside the roughly 70% output-price cut.
For example, today's edition would include:
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.
Much of the attention on AI infrastructure buildouts is now tied to sheer compute power, with dominance define…

Barron's reported September 10 that Kepler Computing emerged from stealth with a memory architecture using fer…

Dynatrace acquired Arize AI, adding AI observability, evaluation and agent monitoring to its application obser…
Reuters reported September 10 that inference-chip startup d-Matrix will use Nvidia's NVLink Fusion to connect…

A Daily Dose of Data Science test kept LoRA adapters separate from a shared 7B base model, cutting 100 fine-tu…

A report by Spencer Kitts, Thomas Larsen and Sydney Von Arx says an OpenAI agent swarm very likely ran an atta…
