Speculative Decoding
LLMs generate one token per step, and each step leaves most of the GPU idle. Speculative decoding fixes the wait without changing the words: a cheap "draft" model guesses several tokens ahead, the real model verifies them all in one parallel pass, and a rejection-sampling rule guarantees the output is statistically identical to the big model alone. Latency drops 2-3×; quality does not move at all.
01.The Problem: A Very Expensive Metronome
Remember how an LLM writes (previous topic, in one plain sentence: it generates one token at a time, reading its whole KV cache each step).
That serial loop is a latency prison.
Every token requires one full forward pass through the model. And each pass — for a 70B model — streams hundreds of megabytes of weights and cache from memory while computing… a single token's worth of arithmetic.
So the GPU is sitting there doing what it does worst:
waiting.
Because decoding is memory-bandwidth-bound (you met that idea in the KV-cache topic), a single step uses a small slice of the chip's compute and then goes idle. The FLOPs are not gone — they are unused. There is a lump of free, wasted arithmetic sitting inside every decode step.
The natural question:
Can we make each expensive step produce more than one token?
You cannot just run the big model twice in parallel on its own future inputs — token 5 needs token 4 to exist. Generation is inherently serial for the verifier.
But you can guess ahead cheaply with something small, and then let the big model check all the guesses at once. Checking is not generation: verifying five candidate positions is exactly the shape of work GPUs were built for — parallel across positions, arithmetic-heavy, riding on that idle compute.
Guess cheap. Verify in parallel. Keep only what the big model would have said.
That is speculative decoding (Leviathan et al., Google, arXiv 2211.17192; independently Chen et al., DeepMind, arXiv 2302.01318).
Speculative Decoding: Draft Cheaply, Verify in Parallel ✅
Speculative Decoding: Draft Cheaply, Verify in Parallel ✅
A cheap drafter proposes γ tokens; the target model verifies all of them in one parallel pass with rejection sampling — output distribution is provably identical to standard sampling, while wall-clock speed comes from batching verification.
Unlock Topic #215: Speculative Decoding
You are viewing a preview. The full in-depth technical walkthrough, worked derivations, and code notebooks for this concept, along with self-assessment quizzes, are available with Pro or Lifetime Access.
Failure modes, high-throughput bottlenecks, and real FAANG implementation decisions.
Interactive system topology diagrams, live parameter simulators, and downloadable SVG charts.
Staff-level multiple-choice quiz questions with instant feedback and answer explanations.
Firebase Google authentication automatically syncs your completed topics and quiz scores.
How clear and actionable was this distributed systems breakdown?