
Alibaba released Qwen3.8-Flash-Next, a cost-efficient AI model.
It matches larger models while using fewer active parameters.
It also offers significant savings in training costs and API pricing.
What happened
Alibaba's Qwen team released Qwen3.8-Flash-Next, a multimodal mixture-of-experts model. It has 125 billion total parameters but only activates 6 billion per token.
Why it matters
The model delivers better results than Qwen3.7-Plus at roughly one-ninth the training cost, according to the Qwen team. It leads the majority of tested tasks against much larger models like DeepSeek-V4-Flash and Claude Opus 4.6.
What to watch
The production version, Qwen3.8-Flash, is priced at $0.16 per million input tokens and $0.47 per million output tokens through QwenCloud. This is about one-twelfth the cost of the flagship Qwen3.8-Max and is seen as pressure on rivals like OpenAI and Anthropic.
Ask the AI about this article →
The release of Qwen3.8-Flash-Next is positioned as an architecture preview for the upcoming Qwen4. A key innovation is the N-gram embedding layer, which stores common word groups in a 'phrase dictionary' that runs in system RAM, not GPU memory. This approach contributes to the model's cost efficiency, with a total of 125 billion parameters but only 6 billion activated per token.
The model's performance is notable in agentic coding and office tasks. It scored 58.7 on DeepSWE and 62.5 on SWE-bench Pro, surpassing both DeepSeek-V4-Flash and Claude Opus 4.6. The gap is wider on office productivity, where it hit 73.9 on CoWorkBench compared to DeepSeek's 45.1, and 55.7 on JobBench. Claude Opus 4.6 only outperforms it on Humanity's Last Exam, which the article notes is an older model from February 2026.
The pricing strategy continues to apply pressure on competitors. Flash-Next costs about one-twelfth as much as the flagship Qwen3.8-Max, with a roughly 12x price gap on tokens. This keeps pressure on OpenAI and Anthropic, especially since OpenAI recently responded with discounts on its own GPT-5.6 models. This trend is good for users but potentially challenging for AI providers that need rapid revenue growth to sustain their investment narratives.
For example, today's edition would include:
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.
AI system scaling has pushed interconnect requirements inside data centers from chips and boards up to racks…

Chinese large-model developer Z.ai says it can now support large-scale inference using roughly 100,000 domesti…

Analyst Ming-Chi Kuo says Nvidia has revived the Rubin CPX AI accelerator with a substantially redesigned arch…

Palantir Technologies stock has posted multi-year gains, including an 11x return over 3 years

Apple has escalated its legal battle against OpenAI, claiming in a new court filing that OpenAI is actively de…

Samsung Electronics has locked up as much as 70% of its memory production capacity under long-term supply agre…
