TOPIC #153Advanced 11 min read

Top-k and Top-p (Nucleus) Sampling: Trimming the Tail Without Killing It

AI
AI & ML Editorial
Report an issue
Key takeawayCore Concept Summary

Left alone, an LLM's word lottery eventually picks something silly — the probability "tail" is huge and full of nonsense. Top-k chops off everything except the best k candidates; nucleus (top-p) keeps only the smallest group of words that adds up to p of the probability. This topic covers why unbounded sampling degenerates, why the fixed cutoff of top-k fails across contexts, and why nucleus became temperature's default companion.

01.The Problem: A Lottery With 100,000 Tickets

Remember temperature (Topic 152): the model scores every possible next word, softmax turns the scores into probabilities, and a draw picks the winner.

Sounds fine. But the vocabulary has 100,000+ tokens.

Even tiny probabilities are nonzero probabilities. And a lottery over 100,000 tickets means the tiny ones keep winning — eventually.

That is where "screwberry" comes from: plausible-looking non-words, and abrupt topic jumps, appear because the tail of a distribution over all valid continuations is enormous and full of nonsense.

Insight

So if the tail is the problem, can we just... cut the tail off?

Holtzman et al. (2019, "The Curious Case of Neural Text Degeneration", arXiv 1904.09751) showed the problem is real and worse than it looks: greedy and wide-sampling both degenerate — repetition loops on one side, incoherent drift on the other. And they found the key structure: human-written text occupies a small, sharply peaked region of the distribution.

Truncation strategies exploit exactly that structure — two of them won the field:

  • Top-k (Fan, Lewis & Daumé, 2018): sort candidates by probability, keep the best k (canonical k=40 for generation), renormalize, sample.
  • Top-p / Nucleus (Holtzman et al., 2019): keep the smallest set of tokens whose cumulative probability reaches p (e.g., 0.9), drop the rest, renormalize.

"Renormalize" just means: after throwing tokens away, the survivors' probabilities no longer add to 1, so scale them back up to 100%.

Two Ways to Chop the Tail ✂️

PRO Architecture Blueprint

Two Ways to Chop the Tail ✂️

Top-k fixes the number of candidates regardless of entropy; nucleus sampling fixes the retained probability mass and lets candidate count float with context uncertainty.

Two Ways to Chop the Tail ✂️
100%
Touchpad: Pinch to zoom • Drag to pan
Rendering visual architecture flowchart...
PRO & LIFETIME CURRICULUM

Unlock Topic #153: Top-k and Top-p (Nucleus) Sampling: Trimming the Tail Without Killing It

You are viewing a preview. The full in-depth technical walkthrough, worked derivations, and code notebooks for this concept, along with self-assessment quizzes, are available with Pro or Lifetime Access.

Production Deep Dive

Failure modes, high-throughput bottlenecks, and real FAANG implementation decisions.

Interactive Blueprints

Interactive system topology diagrams, live parameter simulators, and downloadable SVG charts.

Knowledge Assessment

Staff-level multiple-choice quiz questions with instant feedback and answer explanations.

Cross-Device Progress Sync

Firebase Google authentication automatically syncs your completed topics and quiz scores.

Rate This Architecture ChapterFeedback & Rating

How clear and actionable was this distributed systems breakdown?