
What happened
DeepSeek launched V4.1-Flash, a 763B-parameter open-weight model with a causal encoder-decoder architecture. Artificial Analysis scored it 40 on its Intelligence Index, above DeepSeek V4 Pro 0813 and below GLM-5.3-Flash.
Why it matters
At $0.30 per 1M input tokens and $1.20 per 1M output, V4.1-Flash costs roughly 7x less per Intelligence Index task than GLM-5.3 ($2.01) and Kimi K3 ($2.00). It uses 8B active parameters for input and 16B for output.
What to watch
The model's high verbosity—89k tokens per task, 62% more than V4 Pro—is the trade-off that determines whether its low per-token price translates to real savings. Watch whether its SSD-streaming deployment results hold at 200+ TPS.
WHO IT HITSCompanies running long-context AI agents and self-hosted open models will see the most direct effect, since the KV cache footprint is up to 1/8 that of V4 Flash. Inference engineers evaluating cost per task should note the benchmark price gap.
Ask the AI about this article →
Summaries like this, in your inbox every morning.
DeepSeek's release of V4.1-Flash follows a year in which the lab kept a lower profile, letting peers like GLM and Kimi lead on open models. The new model marks a deliberate architectural break from prior DeepSeek generations, introducing a causal encoder-decoder design that separates input and output computation. As Sebastian Raschka noted, the change is substantial enough that DeepSeek "should have called it DeepSeek V5." The naming choice—only a 0.1 version bump—may understate the shift.
The benchmark picture is mixed. V4.1-Flash sits behind some open models on the Artificial Analysis Intelligence Index, but the body argues this reflects benchmarks that don't yet capture what DeepSeek is optimizing for: efficient use of context. The model's reported KV cache footprint is up to 1/8 that of V4 Flash, and local deployment reports show it running on consumerish hardware at 200+ TPS. This suggests DeepSeek is prioritizing servability—the ability to run on offload pipelines and quantized KV stacks—as much as raw benchmark scores.
The stakes hinge on whether the model's verbosity undermines its low per-token pricing. Artificial Analysis measured 89k tokens per Intelligence Index task, 62% more than V4 Pro 0813. If that verbosity holds in real agent workloads, the total cost advantage may narrow. For inference engineers and teams running long-running agents, the question is whether V4.1-Flash's architecture translates to cheaper production deployments, not just cheap benchmark runs.
For example, today's edition would include:
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.
Dynatrace acquired Arize AI, adding AI observability, evaluation and agent monitoring to its application obser…
A Daily Dose of Data Science test kept LoRA adapters separate from a shared 7B base model, cutting 100 fine-tu…

A report by Spencer Kitts, Thomas Larsen and Sydney Von Arx says an OpenAI agent swarm very likely ran an atta…

Simon Willison wrote that many people, himself included, have gone through an existential crisis when a coding…

Stephen Aarons, a New Mexico defense lawyer of over 40 years, was held in direct contempt and fined $5,000 for…

Perplexity cofounder and Chief Strategy Officer Johnny Ho said GPT‑6 Astra can craft communications, edit real…
