AIToday
Large Language ModelsImage GenerationSiliconANGLE AIPublished: Aug 22, 2026, 06:01 JST3 min read

DeepSeek launches multimodal model matching Anthropic's Opus 4.8

DeepSeek launches multimodal model matching Anthropic's Opus 4.8

Key takeaway

  • DeepSeek launched a new multimodal model that outperforms Anthropic's Opus 4.8 on visual benchmarks.

  • The model, V4 Flash Vision Exp, scored 10% higher on two image tasks and beat Opus 4.8 on ALE and ZeroBench.

  • It is currently available only through paid developer access.

3 Key Points

  1. What happened

    DeepSeek unveiled V4 Flash Vision Exp, a multimodal model derived from its V4 Flash algorithm released in April. The new model beat Anthropic's Opus 4.8 on two visual benchmarks—ALE, which contains more than 1,000 multi-step tasks requiring code writing and media interpretation, and ZeroBench, which has 100 image analysis tasks designed to challenge frontier LLMs. V4 Flash Vision Exp outperformed V4 Flash across all seven text-based benchmarks except Cybergym, which evaluates software vulnerability detection.

  2. Why it matters

    The debut signals DeepSeek's competitive positioning in multimodal AI at a time when enterprises increasingly require models that handle both text and images. V4 Flash Vision Exp's stronger performance on image analysis—scoring more than 10% higher on two visual benchmarks—shows the company is closing gaps with established competitors. The model's foundation, V4 Flash, uses a mixture-of-experts architecture with 284 billion parameters that activates only the neural network best suited to each prompt, reducing hardware demands compared to activating the entire model.

  3. What to watch

    V4 Flash Vision Exp is currently available only via DeepSeek's paid developer platform, though the company may release a free version later, consistent with its practice of open-sourcing earlier models. DeepSeek has not disclosed V4 Flash Vision Exp's architecture details. V4 Flash was trained on 32 trillion tokens of data and uses compression techniques (HCA and CSA) that reduce computing power needed to process 1 million-token prompts by 73%.

Ask the AI about this article →

Context & Analysis

DeepSeek's V4 Flash Vision Exp represents an incremental but substantive advance in the company's multimodal capabilities. The model is built on V4 Flash, released in April, which the company has optimized specifically for visual understanding while maintaining text performance. What distinguishes this release is DeepSeek's demonstration of competitive parity with Anthropic's Opus 4.8, a widely used enterprise model, on two concrete image-understanding benchmarks—ALE and ZeroBench. This matters because visual reasoning remains a key differentiator in enterprise AI applications, where models must navigate interfaces, parse diagrams, and interpret screenshots.

The architecture beneath V4 Flash offers insight into DeepSeek's efficiency strategy. By training on 32 trillion tokens and employing a mixture-of-experts design that activates only necessary components per prompt, the model reduces computational overhead relative to dense architectures. The company's use of KV cache compression techniques—reducing computing power for 1 million-token processing by 73%—suggests DeepSeek is competing not just on accuracy but on cost and inference speed, factors that matter to developers and enterprises operating at scale.

The current availability model—paid developer platform only—departs from DeepSeek's stated practice of open-sourcing earlier models, though the company has indicated a free release may follow. This gatekeeping may reflect commercial timing or a desire to gather production feedback before broader release.

FAQ

How does V4 Flash Vision Exp's performance compare to its predecessor?
V4 Flash Vision Exp outperformed V4 Flash across all seven text-based benchmarks except Cybergym. On image analysis, it scored more than 10% higher on two of four visual benchmarks tested.
When will V4 Flash Vision Exp be available for free?
The model is currently available only via DeepSeek's paid developer platform. The company may release a free version later, given its history of open-sourcing earlier models, but no specific date has been announced.
What is V4 Flash's underlying architecture?
V4 Flash is a mixture-of-experts model with 284 billion parameters comprising multiple neural networks of 13 billion parameters each. When a user enters a prompt, the model activates only the neural network best suited to generate an answer, using significantly less hardware than activating the entire LLM.
SiliconANGLE AIRead Original Article

Get the latest Large Language Models news every morning

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytime

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleAMD pauses as Nvidia earnings loom over AI chip race

The AI news that matters, in one minute each morning.

Sign up free