AIToday
Large Language ModelsTHE DECODERPublished: Aug 9, 2026, 22:00 JST

Google converts Gemma 4 to diffusion model, halves training with 50% speed gain

Google converts Gemma 4 to diffusion model, halves training with 50% speed gain

3 Key Points

  1. What happened

    Google DeepMind built DiffusionGemma by retrofitting its existing Gemma-4-26B-A4B model into a diffusion model using less than ten percent of the original training token budget. The two-stage process combines reinforcement learning and sampler distillation (SD·RL), raising quality on reasoning benchmarks by an average of ten points while nearly quadrupling tokens per compute step.

  2. Why it matters

    DiffusionGemma delivers several times the output speed of Gemma 4 and previous diffusion models while maintaining comparable accuracy—it outputs text at about 1,500 tokens per second and produces answers about 50 percent shorter. Unlike standard language models that commit to an answer sequentially, it develops reasoning and output in parallel, allowing it to correct mistakes before finalization; it solves close to 85 percent of Sudoku puzzles correctly after minimal fine-tuning, where the base model fails entirely.

  3. What to watch

    DiffusionGemma trails Gemma 4 on absolute quality benchmarks and occasionally produces repetition loops. The speed advantage holds mainly for single-user scenarios; once about 32 concurrent requests hit the model, standard language models catch up on throughput. Google explicitly calls it experimental and released it under Apache 2.0 license on Hugging Face to accelerate research on text diffusion.

Not sure about something? Ask the AI

Questions and answers are published on this page.

Summaries like this, in your inbox every morning.

Context & Analysis

Google's approach to DiffusionGemma sidesteps the conventional assumption that building a new model requires training from scratch. By retrofitting an existing, well-trained model (Gemma 4) into a diffusion architecture, the team achieved substantial efficiency gains—using less than ten percent of the original training tokens while nearly quadrupling tokens per compute step. The two-stage training strategy merges reinforcement learning (which typically boosts quality) with sampler distillation (which reduces compute steps), creating a unified process that raises reasoning benchmark scores by an average of ten points.

The model's bidirectional reasoning capability—developing answer and reasoning in parallel rather than committing to outputs sequentially—addresses a fundamental constraint of autoregressive language models. This enables self-correction before finalization, as demonstrated by its ability to solve close to 85 percent of Sudoku puzzles after minimal fine-tuning, whereas the base model fails entirely. Structured outputs like JSON or code repairs complete in just two to three refinement steps because the input already constrains most tokens. However, DiffusionGemma's practical advantages come with trade-offs: it trails Gemma 4 on absolute quality benchmarks, occasionally produces repetition loops due to aggressively reduced compute steps, and maintains its speed edge primarily in single-user settings (once about 32 concurrent requests arrive, standard language models catch up on throughput).

Google's explicit framing of DiffusionGemma as experimental reflects this balance. The release aims to accelerate research on text diffusion and provide a foundation for community adaptations—already evident in its adoption by Interfaze for multilingual speech recognition and in radiology report research. The model's architecture, training data, and settings were carried over from Gemma 4, which Google acknowledges are not necessarily ideal for diffusion, suggesting room for specialized optimization.

FAQ
How much faster is DiffusionGemma than Gemma 4?
DiffusionGemma outputs text at about 1,500 tokens per second and produces answers about 50 percent shorter, delivering several times the output speed of Gemma 4 and previous diffusion models. However, the speed advantage holds mainly for single-user scenarios; once about 32 concurrent requests hit the model, standard language models catch up on throughput.
How was DiffusionGemma created?
Google converted the existing Gemma-4-26B-A4B model into a diffusion model using two training stages—first learning to reconstruct noisy text blocks, then applying reinforcement learning combined with sampler distillation (SD·RL). The entire process used less than ten percent of the original training token budget.
Where is DiffusionGemma available?
Google released DiffusionGemma under an Apache 2.0 license on Hugging Face. It is already being used by startup Interfaze for multilingual speech recognition and in a research project on interactive radiology report generation.

Get the latest Large Language Models news every morning

For example, today's edition would include:

  • AWS CloudWatch Omni now generally available with 17 built-in evaluatorsSiliconANGLE AI · 3h ago
  • Google Vids gets Gemini Omni 1.1 Flash, 1080p videoAI Watch (Impress) · 3h ago
  • OpenAI paper: AI can't say "I don't know"Qiita 機械学習 · 3h ago

AI-summarized, only the topics you pick: one digest a day via Email, LINE, or Slack.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.

Questions and answers are published on this page.

Related Articles

Next articleCost shock drives surge in AI model routers for enterprises