
Modal released new DFlash draft models that accelerate AI inference by denoising entire blocks of tokens in parallel rather than one at a time, and by leveraging the target model's internal context representations.
This overcomes a historical 2–3× speedup ceiling in speculative decoding; Qwen 3.5 122B-A10B now reaches over 1000 tokens/sec, up from 250 without speculation.
What happened
Modal released DFlash draft models for Qwen AI models that use block diffusion instead of traditional sequential token generation. Running Qwen 3.5 122B-A10B with these drafters reached over 1000 tokens/sec on a B200, compared with 250 tokens/sec without speculation.
Why it matters
Speculative decoding has been capped around 2–3× speedup because the small draft model that proposes tokens became the bottleneck. DFlash breaks past this by denoising an entire block of tokens in parallel, and by using the target model's own internal representations to guide drafting, which raises acceptance length from a baseline of 3 to over 9.
What to watch
The DFlash draft models are already integrated with vLLM, SGLang, and Transformers, with draft models available on HuggingFace for Qwen and several other model families. Acceptance length maps nearly linearly to speedup—at length 8, Qwen 3.5 27B achieved 5.62× speedup on one B200.
Ask the AI about this article →
For example, today's edition would include:
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.
Google researchers introduced WikiSkill, a framework that gives AI agents a persistent memory of past mistakes…

A YouGov survey of 1,250 U.S

Numerous large language models (LLMs, AI that understands and generates text) from big names like Meta and Goo…

On Aug. 24, Virtuals Protocol launched a new suite of AI agent creation and ownership tools on Solana, enablin…

Meta Platforms released Muse Glimmer, a 30-billion-parameter AI model with open weights, on Aug

AI is now helping solve longstanding math problems, including the so-called Jacobian conjecture, and top mathe…
