
What happened
Google DeepMind built DiffusionGemma by retrofitting its existing Gemma-4-26B-A4B model into a diffusion model using less than ten percent of the original training token budget. The two-stage process combines reinforcement learning and sampler distillation (SD·RL), raising quality on reasoning benchmarks by an average of ten points while nearly quadrupling tokens per compute step.
Why it matters
DiffusionGemma delivers several times the output speed of Gemma 4 and previous diffusion models while maintaining comparable accuracy—it outputs text at about 1,500 tokens per second and produces answers about 50 percent shorter. Unlike standard language models that commit to an answer sequentially, it develops reasoning and output in parallel, allowing it to correct mistakes before finalization; it solves close to 85 percent of Sudoku puzzles correctly after minimal fine-tuning, where the base model fails entirely.
What to watch
DiffusionGemma trails Gemma 4 on absolute quality benchmarks and occasionally produces repetition loops. The speed advantage holds mainly for single-user scenarios; once about 32 concurrent requests hit the model, standard language models catch up on throughput. Google explicitly calls it experimental and released it under Apache 2.0 license on Hugging Face to accelerate research on text diffusion.
Summaries like this, in your inbox every morning.
Google's approach to DiffusionGemma sidesteps the conventional assumption that building a new model requires training from scratch. By retrofitting an existing, well-trained model (Gemma 4) into a diffusion architecture, the team achieved substantial efficiency gains—using less than ten percent of the original training tokens while nearly quadrupling tokens per compute step. The two-stage training strategy merges reinforcement learning (which typically boosts quality) with sampler distillation (which reduces compute steps), creating a unified process that raises reasoning benchmark scores by an average of ten points.
The model's bidirectional reasoning capability—developing answer and reasoning in parallel rather than committing to outputs sequentially—addresses a fundamental constraint of autoregressive language models. This enables self-correction before finalization, as demonstrated by its ability to solve close to 85 percent of Sudoku puzzles after minimal fine-tuning, whereas the base model fails entirely. Structured outputs like JSON or code repairs complete in just two to three refinement steps because the input already constrains most tokens. However, DiffusionGemma's practical advantages come with trade-offs: it trails Gemma 4 on absolute quality benchmarks, occasionally produces repetition loops due to aggressively reduced compute steps, and maintains its speed edge primarily in single-user settings (once about 32 concurrent requests arrive, standard language models catch up on throughput).
Google's explicit framing of DiffusionGemma as experimental reflects this balance. The release aims to accelerate research on text diffusion and provide a foundation for community adaptations—already evident in its adoption by Interfaze for multilingual speech recognition and in radiology report research. The model's architecture, training data, and settings were carried over from Gemma 4, which Google acknowledges are not necessarily ideal for diffusion, suggesting room for specialized optimization.
For example, today's edition would include:
AI-summarized, only the topics you pick: one digest a day via Email, LINE, or Slack.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.
A security researcher identified a zero-day vulnerability in the macOS version of Meta's Muse AI agent

Adobe announced in September 2026 it is expanding its connector for Google Gemini and its plugin for Claude, s…

Elon Musk reposted Gizmodo senior editor Ray Wong's criticism that Meta's Muse AI agent allegedly gave a Faceb…

A buyer named Usman arrived around 9:15 for an MX Keys Mini pickup, waited and messaged, then left angry at 9:…
Hitachi began offering its "Vulnerability/Patch Discovery & Analysis Report Service" on September 28, and JPX…

OpenRouter launched Jev Router, which feeds each prompt to Jev to judge difficulty, then assigns easy tasks to…
