
What happened
China's DeepSeek announced DeepSeek-V4.1-Flash on September 10, saying the new model's architecture and training methods let it exceed parts of DeepSeek-V4-Pro, GPT-5.6 Sol and Claude Opus 5.
Why it matters
The model is a 552B-parameter MoE design that cuts the KV cache, so off-peak input pricing falls to $0.15 (cache miss) from $0.22 — a same-day reversal of the August 13 price hike.
What to watch
Whether the claimed edge over DeepSeek-V4-Pro and the rival models holds up, since DeepSeek's own benchmark chart is the source. Watch peak-hour pricing, which is double the off-peak rates.
WHO IT HITSDevelopers and companies paying per-token API fees for text-generating AI models are the ones who feel this, since the biggest listed drop is on input costs at off-peak hours.
Ask the AI about this article →
Summaries like this, in your inbox every morning.
DeepSeek's announcement pairs a performance claim with a price cut on the same day. The company says DeepSeek-V4.1-Flash partly surpasses its own flagship DeepSeek-V4-Pro, as well as OpenAI's GPT-5.6 Sol and Anthropic's Claude Opus 5 — a comparison it presents through its own benchmark chart.
The cost story is tied to how the model is built. As an MoE model with 552B total parameters, it activates only some of its processing roles depending on the input, and a new architecture reduces the KV cache, the mechanism that retains and reuses past computations. DeepSeek says this is what lowers input pricing in particular.
The timing is notable: the cut lands the same day as the launch and reverses an August 13 API price increase, so off-peak input costs drop to $0.15 (cache miss) from $0.22. Whether that combination of claims and pricing shifts developer choices is likely to hinge on independent testing of the performance gap, and on how many workloads can realistically run at off-peak hours, where the published rates apply.
For example, today's edition would include:
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.
Michael Burry, famous for The Big Short, is short Nvidia, Palantir and Tesla, and in his Substack newsletter s…

Investors have three creative routes to Anthropic exposure before its expected IPO: buying Alphabet, Amazon, o…

DeepSeek launched V4.1-Flash, a 763B-parameter open-weight model with a causal encoder-decoder architecture

Much of the attention on AI infrastructure buildouts is now tied to sheer compute power, with dominance define…

Barron's reported September 10 that Kepler Computing emerged from stealth with a memory architecture using fer…

Dynatrace acquired Arize AI, adding AI observability, evaluation and agent monitoring to its application obser…