
What happened
Suwesh Prasad Sah's Sept 28 study measured Gemma 4 on one 24GB NVIDIA A10G across 360 requests, finding 1.91–2.19x output throughput with Multi-Token Prediction versus standard generation.
Why it matters
The speedup came from cutting repeated per-token work, not from faster single operations, so the same hardware can deliver roughly double the output rate on Gemma 4.
WHO IT HITSTeams serving Gemma 4 models on their own GPUs — particularly those running vLLM 0.24.0 on 24GB-class hardware — can test whether adding the assistant drafter model roughly doubles generation speed without upgrading their GPU.
Summaries like this, in your inbox every morning.
The study by Suwesh Prasad Sah, published Sept 28 and cited here as a preprint technical report, does not address a new Gemma 4 release. It isolates one variable — whether an auxiliary drafter model that proposes the next tokens is used — while holding the target model and the hardware fixed.
Gemma 4's drafter is not a small standalone language model placed alongside the target; it draws on the target's activations and KV cache. When both proposed candidates are accepted, the target adds one more token, advancing up to three tokens per verification round. That is why the mechanism can pay off even though the drafter itself adds work.
The study's measurements cover 360 requests in total, and output validation succeeded for 172 of 180 MTP trials and 177 of 180 standard trials. The remaining 11 cases were kept in the speed totals rather than discarded, and when only correctly completed outputs were counted, the roughly twofold reduction on reasoning problems still held. The reported figures come from the preprint and were arithmetically rechecked by the article's author, who did not rerun the inference.
Pick your industry and the AI tools you use, and get news related to your work every day.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.
Valuence's "Generative AI User Trends Survey" (June 2024–May 2026) tracked six services including ChatGPT, Gem…

OpenAI began offering a ChatGPT feature that accepts uploaded audio files and can transcribe them, summarize t…

On a self-built 76-question Jev-format set, six trained 3B–9B open-weight models (Imajev-4B, Clef-Flash 9B, Je…

At Gemini at Work, Google introduced the Gemini agent, built into Gemini Enterprise, which gathers information…

Google released Google AI Edge Foresight, a free macOS Labs app that uses EmbeddingGemma 2 and Gemma 4 to tran…

The Association for Human Mathematics said OpenAI's release of 722 AI-generated "mathematical results" is not…
