AIToday
Large Language ModelsTHE DECODERPublished: Aug 22, 2026, 06:00 JST2 min read

DeepSeek releases vision model matching Opus 4.8 on agent tasks

DeepSeek releases vision model matching Opus 4.8 on agent tasks

Key takeaway

  • DeepSeek has released an experimental vision model that combines image understanding with its V4-Flash text capabilities.

  • On internal benchmarks, it nearly matches Anthropic's Opus 4.8 on agent tasks.

  • The model works with OpenAI and Anthropic APIs and handles multiple image formats up to 8,192 pixels per side.

3 Key Points

  1. What happened

    Chinese AI company DeepSeek released V4-Flash-Vision-Exp, an experimental multimodal model that adds image understanding to its text capabilities. On DeepSeek's internal multimodal agent benchmarks, the model scores close to Opus 4.8.

  2. Why it matters

    DeepSeek-V4-Flash-Vision-Exp extends the base text model while preserving its reasoning and world knowledge performance. The vision variant is designed for agent-based applications—it can describe images, extract text from screenshots, analyze diagrams, and work with OpenAI's Chat Completions and Responses APIs as well as Anthropic's Messages endpoint.

  3. What to watch

    Developers can send images three ways—Base64 encoding, publicly accessible URLs (up to 32 MiB), or DeepSeek's free Files API (64 MiB limit). A single request supports up to 600 images; each image costs at most 384 tokens, and pricing follows V4-Flash rates. DeepSeek also released version 0.1.1 of its Harness framework, which supports the new model out of the box.

Ask the AI about this article →

Context & Analysis

DeepSeek-V4-Flash-Vision-Exp represents the company's push into multimodal agent workflows, extending its existing V4-Flash model with vision capabilities while maintaining the base model's text performance in reasoning and world knowledge. The model's near-parity with Opus 4.8 on DeepSeek's internal multimodal agent benchmarks suggests competitive positioning in the visual reasoning space. The compatibility with both OpenAI's and Anthropic's APIs—Chat Completions, Responses APIs, and the Messages endpoint—signals that DeepSeek is designed for integration into existing workflows rather than as a standalone alternative. The release of version 0.1.1 of DeepSeek's Harness framework with out-of-the-box support for the new model underscores the company's investment in developer tooling and ecosystem development.

FAQ

How do you send images to DeepSeek-V4-Flash-Vision-Exp?
Developers can embed images directly with Base64 encoding, point to publicly accessible URLs (up to 32 MiB), or use DeepSeek's new free Files API, which allows uploading a file once and referencing it across multiple requests with a size limit of 64 MiB.
How many images can a single request include?
A single request can include up to 600 images. The maximum edge length is 8,192 pixels per side, but that drops to 4,096 pixels once a request contains 15 or more images.
What is the token cost per image?
Each image costs at most 384 tokens regardless of original resolution. An optional "detail" field can downscale images to 512 x 512 pixels to save tokens when fine visual detail is not needed.

Get the latest Large Language Models news every morning

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytime

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleEnterprise AI fails without fixing legacy systems first

The AI news that matters, in one minute each morning.

Sign up free