
DeepSeek unveiled V4 in two versions: V4-Pro for coding and complex tasks, and V4-Flash for faster, cheaper operation. Both can process 1 million tokens and offer reasoning modes. V4-Pro costs $1.74 per million input tokens and $3.48 per million output tokens; V4-Flash costs about $0.14 per million input tokens and about $0.28 per million output tokens.
V4 uses a redesigned attention mechanism (the part of an AI that weighs which parts of a prompt matter most) that compresses older information instead of treating all earlier text equally. In a 1-million-token context, V4-Pro uses only 27% of the computing power and 10% of the memory required by its previous model V3.2; V4-Flash uses just 10% of the computing power and 7% of the memory.
V4-Pro matches the performance of Anthropic's Claude-Opus-4.6, OpenAI's GPT-5.4, and Google's Gemini-3.1 on major benchmarks, and exceeds other open-source models such as Alibaba's Qwen-3.5 and Z.ai's GLM-5.1 on coding, math, and STEM problems. More than 90% of 85 experienced developers surveyed included V4-Pro in their top model choices for coding tasks.
V4 is DeepSeek's first model optimized for domestic Chinese chips, specifically Huawei's Ascend. Huawei said its Ascend supernode products based on the Ascend 950 series would support DeepSeek V4, allowing users to run modified versions on Huawei chips.
Ask the AI about this article →
For example, today's edition would include:
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.
Anthropic updated the system prompt for Claude 5.1, adding a strict ban on reproducing song lyrics, poems, or…

The Pentagon added OpenAI's ChatGPT Mil and xAI's Grok for Government to its AI platform GenAI.mil, which prev…

Amazon Web Services (AWS) has started offering “AWS Cloud Quest 2.0,” a new version of its online game that te…

A job seeker named Christopher, after five unanswered AI interviews with recruiter 'Riley' from IT firm Everfo…

Sandisk says its NAND-based High Bandwidth Flash (HBF) technology can match HBM bandwidth while providing eigh…

World Labs unveiled Atlas, an omni-model trained on text, images, video, and 3D data that anchors every input…
