TOPIC #128Intermediate 15 min read

WordPiece: Merging by Likelihood, Not Frequency

AI
AI & ML Editorial
Report an issue
Key takeawayCore Concept Summary

WordPiece runs the same greedy merge loop as BPE but hires a stricter referee: a merge only earns a vocabulary slot if the pair appears meaningfully more often than its halves would predict by chance. It is the tokenizer behind BERT's famous 30,522-entry vocabulary.

WordPiece Merge Selection 📊

WordPiece shares BPE's greedy framework but swaps the criterion: it keeps the merge that most increases the probability of the training data, and can stop early when no merge clears the likelihood threshold.

WordPiece Merge Selection 📊
100%
Touchpad: Pinch to zoom • Drag to pan
Rendering visual architecture flowchart...

01.The Problem: Raw Counts Can Glue the Wrong Partners

Recall the last concept in one plain sentence: BPE builds a vocabulary by repeatedly gluing the most frequent neighboring pair of symbols.

Counting feels neutral. But counts can be misleading.

Suppose pieces "t" and "he" keep appearing next to each other.

Is that because "the" is a real, meaningful unit?

Or just because "t" is everywhere anyway — it is the most common letter in English — and "he" is common too? Two popular kids standing next to each other at lunch tell you nothing about friendship.

Raw frequency cannot tell these two stories apart. A pair made of two super-common halves racks up a big count for free, and BPE — greedy counter that it is — may hand that coincidence a precious vocabulary slot, while a rarer-but-genuinely-fused pair waits its turn.

So the question becomes:

Insight

Can we score a merge by how much the pairing beats coincidence, instead of by sheer count?

Google's speech group asked exactly that years before BPE conquered NLP. The answer is WordPiece.

02.The Idea in Plain Words: Merges Must Earn Their Slot

WordPiece was introduced in the Google speech-recognition group (Schuster & Nakajima, 2012) for building token sets out of huge multilingual transcriptions. It became famous as the tokenizer of BERT — the 30,522-entry WordPiece vocabulary of bert-base-uncased is one of the most shipped vocabularies in NLP history.

The algorithm is a near-copy of BPE's greedy loop:

Insight

Start from characters. Repeatedly combine adjacent symbols into new vocabulary entries until a target size is reached.

The single difference is the selection criterion — and that difference changes which merges survive:

  • BPE: merge the most frequent adjacent pair.
  • WordPiece: merge the pair that most increases the likelihood of the training corpus — roughly, keep a merge only if the combined piece occurs often enough relative to how often its halves occur separately.

Same loop, different referee.

Why did a speech group invent this? Their problem was text-free: they had enormous multilingual transcriptions and needed a symbol set that covered many languages without hand-building alphabets per language. The same greedy-merge machinery that assembles subword vocabularies assembles sub-phoneme units — and the "does this pairing beat chance?" filter is exactly what you want when languages share fragments for accident versus when they share real structure.

Say "likelihood" out loud and unpack it once:

Insight

The likelihood of the corpus = how probable the whole training text looks if your vocabulary (and its splitting habits) were a tiny theory of language.

A merge increases that likelihood when the fused symbol genuinely compresses the data: the text becomes more predictable under the new vocabulary. A merge that merely joins two ubiquitous halves changes nothing — the "theory" gains a redundant rule.

A house-building analogy: BPE nails together whatever two bricks touched most often. WordPiece inspects each joint first and asks: is this connection real, or are both bricks just lying around everywhere? Only real joints get welded.

03.The Likelihood Criterion, Intuited — With Tiny Numbers

WordPiece scores every candidate merge of pieces A and B with a ratio of the form:

score = P(AB) / (P(A) x P(B))

where probabilities are corpus frequencies. Read it out loud:

Insight

observed-togetherness ÷ expected-togetherness-by-chance

A merge is added only when the pair appears more often than chance from its parts — i.e., when treating AB as one symbol makes the corpus more predictable. The merge with the top score is applied, then candidates are re-scored.

Now tiny numbers. Say a toy corpus gives:

Case 1 — "t" + "he" → "the"

  • P(t) = 0.020, P(he) = 0.030 → chance expectation = 0.020 x 0.030 = 0.0006
  • Observed P(the) = 0.015
  • Score = 0.015 / 0.0006 = 25 → the pair is 25x more common than coincidence predicts. Real fusion. Weld it.

Case 2 — "ing" + "s" → "ings"

  • P(ing) = 0.090, P(s) = 0.050 → chance = 0.0045
  • Observed P(ings) = 0.0050
  • Score = 0.0050 / 0.0045 ≈ 1.1 → barely above chance. The raw count is big only because both halves are common on their own. WordPiece shrugs and moves on.

Two pairs, similar counts, opposite verdicts. That is the whole difference between the two algorithms.

And notice what happens after each accepted weld: pieces are re-counted, and brand-new candidate pairs appear (the can now pair with r to try for ther). The loop is greedy — one best merge, then re-scan, forever — exactly like BPE, only with the stricter referee at each step.

The classic worked example from the original paper uses Chinese:

  • Tokens "辛" (hard) and "苦" (bitter) exist separately because each is independently frequent.
  • Their concatenation "辛苦" (thank-you-for-your-hard-work) is also kept, because as a phrase it occurs far above the chance baseline — even though "辛" alone is relatively rare.

Frequency alone would make different choices here: WordPiece is explicitly protecting against merging pairs that only look common because both halves are common on their own (exactly the failure mode of raw BPE on words like "the" split as "t" + "he").

One more consequence: training stops at the target vocabulary size or when the best score drops below a threshold.

So WordPiece can walk away early, and its vocabularies can end up slightly smaller and more conservative than BPE's. BPE always fills the menu; WordPiece stops serving dishes nobody actually craves together.

04.Visual Intuition: The Coincidence Detector

Picture two coworkers, Dana and Priya. You keep seeing them together in the elevator.

Dana takes that elevator 50 times a week. Priya takes it 50 times a week. Same building, same floor, one elevator.

Of course you see them together. Zero information.

Now imagine Priya rarely leaves her 9th-floor lab, yet is always next to Dana. Now you suspect something: they work together.

WordPiece is that suspicion engine, applied to symbols:

code
        BPE referee              WordPiece referee
   ┌──────────────────┐      ┌────────────────────────┐
   │  count(AB) = ?   │      │  count(AB)             │
   │                  │      │  ───────────────────   │
   │  biggest number  │      │  count(A) x count(B)   │
   │  wins the slot   │      │                        │
   │                  │      │  ratio > 1 by a lot?   │
   └──────────────────┘      │  then it is real       │
                             └────────────────────────┘

   chance baseline:         ⚖ balanced scale ⚖
   P(A)·P(B)  ←—————————→  P(AB)
   tips only when AB genuinely outweighs coincidence

Frequency asks "how much?". Likelihood asks "how much MORE than expected?" — and that second question is the difference between correlation and coincidence.

05.The ## Continuation Marker and the Fate of Capital Letters

WordPiece piece-level conventions differ visually from BPE:

  • The first piece of a word is bare: token
  • Every continuation piece carries a ## prefix: ##ization, ##ness
  • A word boundary is thus written into the token string itself — unlike byte-level BPE, where the space hides inside tokens like Ġis.

Why care? Because it makes splits human-readable. Glance at hydro ##ele ##ctri ##city and the word assembly is obvious. Debugging tokenizers at 2 a.m. is much kinder when the seams are labeled.

The canonical BERT vocabularies come in two flavors:

  • Uncased: text is lower-cased before tokenization. Case information is destroyed — fine for classification, bad for generating exact strings back.
  • Cased: 30,522 WordPiece entries preserving upper/lowercase distinctions — at the cost of fragmenting capitalized rare words more.

Watch what uncasing does to a name. Input "UNICEF" arrives, the pipeline shouts it down to "unicef", the tokenizer slices uni + ##cef, and the model outputs... "unicef". The organization's name is understood but its spelling is gone — the capital-letter evidence was thrown away before the model ever woke up. That is why extractive tasks (names, codes, capitalized entities) use cased variants, and why generative pipelines trained on uncased BERT-style vocabularies can never promise exact-case output.

And different families mark the same idea differently: SentencePiece (ALBERT, T5, LLaMA-1/2) uses a visible underscore character (U+2581, "▁") as a prefix space instead of ##. Mixing these conventions across checkpoints is a classic source of silent tokenization drift — your vectors are fine, your units quietly are not.

python— Inspecting WordPiece behavior in HuggingFace
from transformers import BertTokenizer

tok = BertTokenizer.from_pretrained("bert-base-uncased")  # vocab size: 30,522
print(tok.tokenize("Hydroelectricity is inexpensive hydropower"))
# ['hydro', '##ele', '##ctri', '##city', 'is', 'in', '##ex', '##pen', '##sive', 'hydro', '##power']

print(len(tok.vocab))          # 30522
print(tok.tokenize("UNICEF"))   # cased models keep 'UNI', '##CEF'; uncased folds to ['uni', '##cef']

06.In Practice: Where WordPiece Lives in 2024-2026

WordPiece never became the dominant tokenizer for generative models — the GPT/LLaMA lineage is BPE through and through. But it remains all over the encoder ecosystem:

  • BERT-class models: the original BERT, BERTweet, many legal/medical/biomedical BERTs, and thousands of sentence-similarity and classification fine-tunes still load their 30k WordPiece vocabulary.
  • DeBERTa-v3 (Microsoft, widely used in 2024–2026 rerankers, e.g., cross-encoders for RAG) ships a ~128k SentencePiece tokenizer with ## continuation markers — a deliberate hybrid, which is why some tokenizers expose a wordpieces_split option.
  • Embedding services built on BERT-family checkpoints must keep WordPiece in the inference binary even when the serving stack is otherwise GPT-native.

Practical guidance, in one paragraph: WordPiece, BPE, and Unigram produce nearly identical IDs on common English text. The differences bite on rare words, code-switching, non-Latin scripts, and exact string reconstruction — and none of it can be changed after training. The referee you pick at menu-writing time is welded into every checkpoint that menu ever serves.

So when you meet a BERT-class model and ask "why 30,522 slots and not 100k?" — the answer is WordPiece's conservative streak: a likelihood threshold plus early stopping, frozen in 2018, still serving billions of requests a day.

Architectural Trade-offs & Production Realities

Architectural Advantages

  • Likelihood criterion avoids spurious merges of pairs that are common only because their halves are common.
  • Early stopping on threshold yields compact, conservative vocabularies (BERT: 30,522 entries).
  • Explicit ## boundary markers make tokenization human-debuggable at a glance.

Trade-offs & Constraints

  • Uncased variants permanently destroy case information needed for generation and entity strings.
  • Slower to train than BPE: re-scoring likelihood ratios over candidates is heavier than counting pairs.
  • Smaller vocabulary than modern BPE rivals fragments long rare words and non-English text more aggressively.
Production Implementation in Big Tech
Google Research / Hugging Face• WordPiece from ASR symbol tables to BERT checkpoints

Google originally used WordPiece to build language-independent symbol sets for multilingual speech recognition. Eight years later, the same algorithm ships the fixed 30,522-token vocabulary inside every bert-base-uncased deployment — recommendation filters, ad relevance classifiers, and RAG rerankers decode requests through it byte-for-byte as trained.

Staff+ Engineering Takeaways

  • WordPiece = BPE framework with a likelihood criterion: a merge is kept only if it increases the probability of the training corpus.
  • Score roughly P(AB) / (P(A) x P(B)); training stops at the size target or when no merge beats a threshold.
  • Continuation pieces are marked with ##, making word boundaries explicit; SentencePiece uses the ▁ underscore convention instead.
  • BERT's canonical vocabulary is 30,522 WordPiece entries, in uncased and cased flavors.
  • WordPiece anchors the encoder ecosystem (BERTs, many rerankers) while the generative ecosystem standardized on BPE.

Topic Knowledge Check

Exercise 1 of 3 • Test your architectural comprehension.

Exercise 1 of 30 answered
1

What is the core algorithmic difference between WordPiece and BPE?

Rate This Architecture ChapterFeedback & Rating

How clear and actionable was this distributed systems breakdown?