TOPIC #215Advanced 14 min read

Speculative Decoding

AI
AI & ML Editorial
Report an issue
Key takeawayCore Concept Summary

LLMs generate one token per step, and each step leaves most of the GPU idle. Speculative decoding fixes the wait without changing the words: a cheap "draft" model guesses several tokens ahead, the real model verifies them all in one parallel pass, and a rejection-sampling rule guarantees the output is statistically identical to the big model alone. Latency drops 2-3×; quality does not move at all.

01.The Problem: A Very Expensive Metronome

Remember how an LLM writes (previous topic, in one plain sentence: it generates one token at a time, reading its whole KV cache each step).

That serial loop is a latency prison.

Every token requires one full forward pass through the model. And each pass — for a 70B model — streams hundreds of megabytes of weights and cache from memory while computing… a single token's worth of arithmetic.

So the GPU is sitting there doing what it does worst:

Insight

waiting.

Because decoding is memory-bandwidth-bound (you met that idea in the KV-cache topic), a single step uses a small slice of the chip's compute and then goes idle. The FLOPs are not gone — they are unused. There is a lump of free, wasted arithmetic sitting inside every decode step.

The natural question:

Insight

Can we make each expensive step produce more than one token?

You cannot just run the big model twice in parallel on its own future inputs — token 5 needs token 4 to exist. Generation is inherently serial for the verifier.

But you can guess ahead cheaply with something small, and then let the big model check all the guesses at once. Checking is not generation: verifying five candidate positions is exactly the shape of work GPUs were built for — parallel across positions, arithmetic-heavy, riding on that idle compute.

Guess cheap. Verify in parallel. Keep only what the big model would have said.

That is speculative decoding (Leviathan et al., Google, arXiv 2211.17192; independently Chen et al., DeepMind, arXiv 2302.01318).

Speculative Decoding: Draft Cheaply, Verify in Parallel ✅

PRO Architecture Blueprint

Speculative Decoding: Draft Cheaply, Verify in Parallel ✅

A cheap drafter proposes γ tokens; the target model verifies all of them in one parallel pass with rejection sampling — output distribution is provably identical to standard sampling, while wall-clock speed comes from batching verification.

Speculative Decoding: Draft Cheaply, Verify in Parallel ✅
100%
Touchpad: Pinch to zoom • Drag to pan
Rendering visual architecture flowchart...
PRO & LIFETIME CURRICULUM

Unlock Topic #215: Speculative Decoding

You are viewing a preview. The full in-depth technical walkthrough, worked derivations, and code notebooks for this concept, along with self-assessment quizzes, are available with Pro or Lifetime Access.

Production Deep Dive

Failure modes, high-throughput bottlenecks, and real FAANG implementation decisions.

Interactive Blueprints

Interactive system topology diagrams, live parameter simulators, and downloadable SVG charts.

Knowledge Assessment

Staff-level multiple-choice quiz questions with instant feedback and answer explanations.

Cross-Device Progress Sync

Firebase Google authentication automatically syncs your completed topics and quiz scores.

Rate This Architecture ChapterFeedback & Rating

How clear and actionable was this distributed systems breakdown?

Related Concepts & Cross-References

Indexed from curriculum