AIToday
Large Language ModelsOpen-Source AIQiita 機械学習Published: Oct 8, 2026, 19:00 JST

Gemma 4 gets 1.91–2.19x faster with Multi-Token Prediction, Sah finds

Gemma 4 gets 1.91–2.19x faster with Multi-Token Prediction, Sah finds

3 Key Points

  1. What happened

    Suwesh Prasad Sah's Sept 28 study measured Gemma 4 on one 24GB NVIDIA A10G across 360 requests, finding 1.91–2.19x output throughput with Multi-Token Prediction versus standard generation.

  2. Why it matters

    The speedup came from cutting repeated per-token work, not from faster single operations, so the same hardware can deliver roughly double the output rate on Gemma 4.

WHO IT HITSTeams serving Gemma 4 models on their own GPUs — particularly those running vLLM 0.24.0 on 24GB-class hardware — can test whether adding the assistant drafter model roughly doubles generation speed without upgrading their GPU.

Not sure about something? Ask the AI

Questions and answers are published on this page.

Summaries like this, in your inbox every morning.

Context & Analysis

The study by Suwesh Prasad Sah, published Sept 28 and cited here as a preprint technical report, does not address a new Gemma 4 release. It isolates one variable — whether an auxiliary drafter model that proposes the next tokens is used — while holding the target model and the hardware fixed.

Gemma 4's drafter is not a small standalone language model placed alongside the target; it draws on the target's activations and KV cache. When both proposed candidates are accepted, the target adds one more token, advancing up to three tokens per verification round. That is why the mechanism can pay off even though the drafter itself adds work.

The study's measurements cover 360 requests in total, and output validation succeeded for 172 of 180 MTP trials and 177 of 180 standard trials. The remaining 11 cases were kept in the speed totals rather than discarded, and when only correctly completed outputs were counted, the roughly twofold reduction on reasoning problems still held. The reported figures come from the preprint and were arithmetically rechecked by the article's author, who did not rerun the inference.

FAQ
How much faster is Multi-Token Prediction on Gemma 4?
Output throughput rose 1.91–2.19x across nine fixed prompts, but first-output time fell only 10.0–14.2%, so the biggest improvement is in tokens generated after the first one.
Why did it get faster if the GPU kernel itself is not quicker?
The representative kernel took 2.78ms normally and 2.79ms with MTP, but repeated CUDA Graph executions per output token dropped 56.4–78.1%, so the same work is called far fewer times.
Does this guarantee the same speedup in production?
No. The 360 measurements used nine fixed prompts repeated 20 times each with concurrency of 1, so performance under real production load must be measured separately.
Qiita 機械学習Read Original Article

AI news that matters for your work, delivered every morning.

Pick your industry and the AI tools you use, and get news related to your work every day.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.

Questions and answers are published on this page.

Related Articles

Next articleNous Research raises $90 million for corporate AI push