Greedy and Beam Search vs Sampling: How Search Became a Decoding Choice
Once you can score each next word, you have to decide how to build a whole sentence. Three philosophies compete: greedy decoding (take the best next word now), beam search (keep the best B partial sentences and prune), and sampling (roll the probability dice). This topic explains all three — why beams ruled 2016-era machine translation, why they fail at open-ended chat, and why search is now returning inside agents.
01.The Problem: One Word at a Time Is Not a Sentence
You already know the pieces: a language model scores every candidate next token, and something must pick one (Topics 152–153).
But generation is sequential. Pick one word, and the next pick sees it. Pick wrong, and the rest of the sentence inherits the damage.
So the real question is not "which word next?" but
What strategy should build the whole sentence?
Three answers, three philosophies:
- Greedy: always take the highest-scoring next word. Deal with consequences later. (There is no later — see below.)
- Beam search: keep several candidate sentences in play at once, score them as wholes, and throw away the losers step by step.
- Sampling: don't optimize at all — draw words from the probability distribution, dice and all.
Which one "wins" changed the shape of the field twice: beams defined machine translation around 2016; sampling defined the chat era after 2022; and in 2024–2026, search is sneaking back inside agents. This topic is that whole arc.
Three Decoder Philosophies 🔀
Three Decoder Philosophies 🔀
Greedy optimizes each step; beam optimizes the sequence over B hypotheses; sampling accepts the distribution as-is — and each was "best" in its decade.
Unlock Topic #154: Greedy and Beam Search vs Sampling: How Search Became a Decoding Choice
You are viewing a preview. The full in-depth technical walkthrough, worked derivations, and code notebooks for this concept, along with self-assessment quizzes, are available with Pro or Lifetime Access.
Failure modes, high-throughput bottlenecks, and real FAANG implementation decisions.
Interactive system topology diagrams, live parameter simulators, and downloadable SVG charts.
Staff-level multiple-choice quiz questions with instant feedback and answer explanations.
Firebase Google authentication automatically syncs your completed topics and quiz scores.
How clear and actionable was this distributed systems breakdown?