
Speculative decoding works by having a small model (10–100x smaller than the target) generate K candidate tokens, then a large model verifies all K tokens in a single forward pass before accepting or rejecting each token. Google uses this in AI Overviews to serve over a billion Search users; Anthropic, Meta, and major inference providers use it to reduce latency at scale.
Same-tokenizer pairs (draft and target models sharing a tokenizer) achieve 1.5–3x speedup, while cross-tokenizer pairs achieve 1.5–1.9x speedup. For example, Llama 3.2 1B as drafter with a larger target model achieved 2.31x speedup, while Llama 3.1 8B achieved 2.08x despite higher token acceptance.
Emerging variants like EAGLE (trains a lightweight head on the target model's hidden states), Medusa (adds multiple prediction heads), and self-speculative decoding (LayerSkip, SWIFT—uses the target model's own early layers as drafter) aim to eliminate the need for a separate draft model or extra training.
Ask the AI about this article →
For example, today's edition would include:
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.
Walmart settled opioid dispensing claims for $50 million

Tim Cook's legacy as Apple CEO is now tied to the company's push into artificial intelligence, according to a…

CrowdStrike is introducing Falcon Guardian, its flagship solution for the AI Detection and Response (AIDR) cat…

AT&T, Dell Technologies, and AMD have announced OTel 2.0, the largest and best-performing open-source model bu…

AT&T's legal department built an in-house center of expertise called Legal Edge, described as an AI-first lega…

John Deere introduced JD, an AI assistant designed to help farmers manage and interpret their farm data, as re…
