AIToday
Large Language ModelsZenn AI/MLPublished: Oct 2, 2026, 22:00 JST

Qwen3.8-Flash-Next Previews Qwen4.x Architecture

Qwen3.8-Flash-Next Previews Qwen4.x Architecture

3 Key Points

  1. What happened

    Qwen released Qwen3.8-Flash-Next on August 27, 2026, calling it a preview of the architecture planned for Qwen4.x.

  2. Why it matters

    The design suggests Qwen's team is applying its own improvements rather than reusing existing structures, which points to independently developed technology rather than reliance on distillation.

  3. What to watch

    The outcome hinges on whether llama.cpp adds full support for the new attention structure before Qwen4.x arrives. Watch the sparse flash attention pull request, #28349.

WHO IT HITSLocal-LLM enthusiasts and developers running models on consumer GPUs are affected most directly, since the architecture aims to cut memory use and speed up output for large models.

Not sure about something? Ask the AI

Questions and answers are published on this page.

Summaries like this, in your inbox every morning.

Context & Analysis

Qwen3.8-Flash-Next arrived as a surprise pre-release, announced on August 26, 2026 and shipped the following day. What drew attention was the stated link to an architecture intended for Qwen4.x, signaling that Qwen's team is willing to test next-generation designs in public before the main release.

The model departs from previous Qwen designs in two notable ways. Qwen Sparse Attention processes 4-token microblocks rather than single tokens, aiming to reduce computation for long contexts. N-gram Embedding keeps static information in main memory instead of graphics memory, which could lower memory pressure for large models.

The author's testing suggests the approach pays off in practice. Even with a very large model, the context memory footprint came in well below what a smaller Qwen model used. It's worth watching whether the supporting software ecosystem catches up before Qwen4.x lands.

FAQ
What makes Qwen3.8-Flash-Next different from earlier Qwen models?
It uses an architecture intended for Qwen4.x, featuring Qwen Sparse Attention that handles 4-token microblocks instead of single tokens, plus N-gram Embedding stored in main memory.
Can I run Qwen3.8-Flash-Next in llama.cpp today?
Yes, but with limitations. The initial architecture support was merged on August 28, 2026, and full support for the new attention structure is still pending.
How does the microblock approach affect long-context performance?
The author observed that read speed stayed steady even beyond 40,000 tokens, suggesting the new attention structure greatly reduces the slowdown seen with standard attention.

AI news that matters for your work, delivered every morning.

Pick your industry and the AI tools you use, and get news related to your work every day.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.

Questions and answers are published on this page.

Related Articles

Next articleMicrosoft expands biomimicry to more than 20 data center sites