AIToday
Large Language ModelsAI Business & IndustryITmedia AI+Published: Sep 10, 2026, 19:01 JST2 min read

DeepSeek-V4.1-Flash beats flagship, cuts API prices same day

DeepSeek-V4.1-Flash beats flagship, cuts API prices same day

3 Key Points

  1. What happened

    China's DeepSeek announced DeepSeek-V4.1-Flash on September 10, saying the new model's architecture and training methods let it exceed parts of DeepSeek-V4-Pro, GPT-5.6 Sol and Claude Opus 5.

  2. Why it matters

    The model is a 552B-parameter MoE design that cuts the KV cache, so off-peak input pricing falls to $0.15 (cache miss) from $0.22 — a same-day reversal of the August 13 price hike.

  3. What to watch

    Whether the claimed edge over DeepSeek-V4-Pro and the rival models holds up, since DeepSeek's own benchmark chart is the source. Watch peak-hour pricing, which is double the off-peak rates.

WHO IT HITSDevelopers and companies paying per-token API fees for text-generating AI models are the ones who feel this, since the biggest listed drop is on input costs at off-peak hours.

Ask the AI about this article →

Summaries like this, in your inbox every morning.

Context & Analysis

DeepSeek's announcement pairs a performance claim with a price cut on the same day. The company says DeepSeek-V4.1-Flash partly surpasses its own flagship DeepSeek-V4-Pro, as well as OpenAI's GPT-5.6 Sol and Anthropic's Claude Opus 5 — a comparison it presents through its own benchmark chart.

The cost story is tied to how the model is built. As an MoE model with 552B total parameters, it activates only some of its processing roles depending on the input, and a new architecture reduces the KV cache, the mechanism that retains and reuses past computations. DeepSeek says this is what lowers input pricing in particular.

The timing is notable: the cut lands the same day as the launch and reverses an August 13 API price increase, so off-peak input costs drop to $0.15 (cache miss) from $0.22. Whether that combination of claims and pricing shifts developer choices is likely to hinge on independent testing of the performance gap, and on how many workloads can realistically run at off-peak hours, where the published rates apply.

FAQ
How much does DeepSeek-V4.1-Flash cost?
Off-peak rates are $0.003 for input with a cache hit and $0.15 with a cache miss, plus $0.6 for output. Peak-hour rates are double those figures.
How does the new pricing compare with the old rates?
The old rates were $0.007, $0.22 and $0.66 respectively for cache-hit input, cache-miss input and output. DeepSeek had raised prices on August 13 before this cut.
What is DeepSeek-V4.1-Flash built on?
It is an MoE model that runs only part of its processing roles depending on the input, with 552B total parameters. A new architecture reduces the KV cache, which stores and reuses past computations.

Get the latest Large Language Models news every morning

For example, today's edition would include:

  • DeepSeek V4.1-Flash: 763B model beats V4 Pro on AA Index 40Latent Space · 37m ago
  • Dynatrace acquires Arize AI as observability shifts to actionSiliconANGLE AI · 6h ago
  • Shared base cuts 100 fine-tunes from 1.5 TB to 19.3 GBDaily Dose of Data Science · 6h ago

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleApple's iPhone Duo Bet Leaves AI Blindspot