
Qwen released Qwen3.8-Flash-Next, an open-weights AI model. It has 125B total parameters but only 6B active.
The model also previews the architecture for Qwen4.
It runs well on a DGX Spark.
What happened
Alibaba's Qwen team released Qwen3.8-Flash-Next, an open-weights multimodal MoE model that also serves as an early preview of the architecture used in Qwen4.
Why it matters
The model has 125B total parameters but only 6B active, which gives it a significant performance boost while keeping computational costs lower. It can run on a DGX Spark using quantized versions.
What to watch
The model is available in multiple quantized sizes, including a 72.5GB UD-IQ1_S and a 78.9GB UD-Q2_K_XL, with the latter producing the author's favorite results so far.
Ask the AI about this article →
This release is notable because it not only adds another open-weights model to Qwen's lineup but also gives an early look at the architecture that will power Qwen4. By using a mixture-of-experts design with only 6B active parameters out of 125B total, the model aims to balance capability with efficiency, fitting on consumer-grade hardware like the DGX Spark.
The author's hands-on testing with quantized versions produced promising results, particularly with the 78.9GB UD-Q2_K_XL variant. This suggests practical usability for developers who want to experiment with a large multimodal model without needing a massive cluster. The choice of quantized models also indicates that the open-weights community is already building tooling around it, which could accelerate adoption.
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytime
Ask AI anything about this article. Q&As are published on this page for other readers too.
Z.ai Co. released the code for GLM-5.3-Flash, an LLM with 320 billion parameters that activates 18 billion per…
Deep Cogito Inc. raised $43 million in a Series A round led by TQ Ventures, with participation from Benchmark…
Beijing introduced China's first dedicated "AI4Chip" policy, extending AI into semiconductor production across…

Many cloud-based AI services, including ChatGPT, use input data for AI training by default, even on paid perso…

Daily Dose of Data Science built an AI workflow using Mistral OCR 4 that reads every chart in a scientific pap…

Meta's Project OT, an internal plan to use AI agents to replace workers, would have reduced headcount by about…
