Vocabulary Size: Coverage vs. Compute in Tokenizer Design
Every tokenizer ships with one number nobody sees in the demo: the vocabulary size — 30k to 256k slots. That single choice sets embedding-table parameters, softmax cost, how many tokens each language burns per sentence, and how much a model costs to serve. Bigger is not free; smaller is not cheap either.
Vocabulary Size Is a Two-Sided Lever ⚖️
Growing the vocabulary buys coverage and shorter sequences, but charges you in fixed parameters, softmax compute, and diluted training signal for rare pieces.
01.The Problem: One Number Nobody Puts on the Box
You have now met the vending machine (tokenization) and its menu-writing algorithm (BPE, WordPiece, Unigram).
Every menu has a length. The tokenizer for BERT carries 30,522 pieces. GPT-4o carries roughly 200,000. Gemini carries about 256,000.
How was that number chosen? And why should you care?
Because the vocabulary size — usually written
V— silently controls four things at once:
- how many parameters the model burns on the first and last layer,
- how expensive every forward pass is,
- how many tokens (money!) each language costs per sentence,
- how well rare pieces are actually learned.
Pick V too small and everything gets chopped into too many pieces — Japanese text can cost 3–6x more tokens than it should, and every one of them is billed.
Pick V too big and you pay a fixed tax in weights and compute on every single step — plus a long tail of vocabulary rows that barely trained at all.
This concept is about that trade. It is a two-sided lever, and once trained, the lever is welded in place.
02.The Idea in Plain Words: V Is the Menu Length
First, where does V actually live inside a Transformer? Almost nowhere — exactly two places, both at the edges of the network:
- Input embedding matrix: shape
V x d. V rows, each a d-dimensional vector. (From the embeddings concept: a token's ID just picks its row.) - Output projection ("unembedding"): another
V x dmatrix, unless tied with the input. Its job is to score every vocabulary item at every step so the softmax — the layer that turns scores into probabilities — can pick the next token.
Everything in the middle (attention heads, MLPs, layers) does not care about V at all.
So the mental model:
V sets the size of the model's front door and back door. The hallway in between is built to a different spec.
Front door: numbers → vectors. Back door: vectors → probabilities over all V candidates.
Let's put real numbers on the doors.
03.A Tiny Worked Example: Counting Doors
Part A — the parameter bill.
LLaMA-3 has V = 128,256 and hidden size d = 4,096.
V x d = 128,256 x 4,096 ≈ 525,000,000
That is ~0.53 billion parameters in the embedding matrix alone. LLaMA-3 unties its output projection, so the back door is another 0.53B.
0.53B + 0.53B ≈ over 1 billion parameters for a model whose total is 8B.
More than an eighth of the entire model is the menu.
Part B — the token bill.
Take the word internationalization.
- Under a small 30k English-leaning vocabulary it might shatter into
inter + na + tion + al + ization→ 5 tickets. - Under a 200k multilingual vocabulary it is plausibly 1 ticket.
Same meaning. Five times the positions. And positions are where everything hurts: attention cost grows quadratically with sequence length, context windows fill up 5x faster, and per-token API pricing multiplies by five.
Same lever, opposite ends: Part A is the fixed cost of a big menu, Part B is the variable cost of a small one.
# LLaMA-3-8B: 128,256 vocab x 4096 hidden, untied embeddings
V, d = 128_256, 4096
print(f"Embedding table: {V*d/1e9:.2f}B params") # 0.53B
print(f"Unembedding: {V*d/1e9:.2f}B params") # 0.53B (untied)
# vs bert-base-uncased: 30,522 x 768 = 0.023B — tiny because encoders
# classify with small heads instead of predicting over the full vocab04.Visual Intuition: A Seesaw, and a Timeline
The whole topic balances on one seesaw:
codeSMALL V BIG V ┌────────────────┐ ┌────────────────┐ │ tiny doors │ │ 0.5B+ params │ ← fixed tax, │ cheap softmax │ │ fat logit layer│ every pass │ │ │ │ │ every rare word│ │ common words │ │ shatters into │ ◄──seesaw──►│ fit in ONE │ │ many tokens │ │ token │ │ (3-6x for │ │ │ │ Japanese!) │ │ long tail of │ └────────────────┘ │ cold rows │ └────────────────┘
And the industry has been visibly sliding right along that lever:
code2018 BERT 30k 2019 GPT-2 50k 2023 GPT-3.5/4 100k 2024 GPT-4o 200k LLaMA-3 128k 2024+ Qwen 152k Gemini ~256k ────────────────────────────────────────► vocab sizes grew ~5–10x in six years, driven by multilingual + code economics
Nobody did this out of greed for big numbers. Each jump paid for itself in shorter sequences.
05.The Analogy: A Kitchen That Pre-Cooks Everything
Carry one picture through the rest of this concept: a restaurant kitchen deciding how many dishes to pre-cook.
The vocabulary is the set of pre-made dishes. A sentence is an order the kitchen assembles by platter-combining.
- Many pre-made dishes (big V): most orders need only two or three platter grabs — service is fast, the counter (context window) empties quickly. But the rent is brutal: every dish needs its own cabinet slot (an embedding row), and the waiter must check every dish before recommending one (softmax over V).
- Few pre-made dishes (small V): rent is cheap. But every order of "internationalization" requires scrambling five ingredients together from basic stock — slower service, longer lines, and the same popular dish arrives at the table in fragments.
- Rare dishes (the long tail): a dish ordered 5 times all season never gets practiced. Its cook (the embedding row) learns almost nothing from those 5 orders — it stays close to raw ingredients, i.e., a near-random vector the model then misuses.
Three bills, one kitchen. Keep them separate:
Rent = fixed parameters + softmax width. Prep time = tokens per text. Practice = occurrences per vocabulary row.
Growing V cuts prep time, raises rent, and dilutes practice.
06.Why AI Cares — The Coverage Side of the Ledger
A vocabulary too small for its training distribution fragments text: common words get split into several pieces, so every sentence costs more tokens. The measured numbers:
- English under a 30k WordPiece vocabulary averages ~1.4–1.6 tokens per word; under a modern 128k–200k byte-level BPE, ~1.2–1.3.
- For morphologically rich or non-Latin languages the gap explodes: Japanese or Chinese under an English-heavy 50k vocabulary can run 3–6x more tokens than under a 128k+ multilingual vocabulary.
- Longer sequences mean: quadratic attention cost, faster context-window exhaustion, slower inference, and higher price per unit of meaning. This is why tokenizer choice was the headline complaint about early English-centric models serving global traffic.
The trend line in section 4 is this ledger in motion: BERT (2018) 30k → GPT-2 (2019) 50k → GPT-3.5/4 (2023) 100k → GPT-4o (2024) 200k, LLaMA-3 (2024) 128k, Qwen 152k, Gemini ~256k — vocabularies grew ~5–10x in six years, driven almost entirely by multilingual and code economics.
07.Why AI Cares — The Cost Side of the Ledger
Compute per step also scales with V. The final projection produces a V-wide logit vector and the softmax cross-entropy sums over all of it. For large V and long contexts, frameworks shard this computation across GPUs (tensor/vocabulary parallelism, e.g., Megatron-LM) precisely because the V x d matmul is the single widest linear layer in the stack.
Why not just ship a 1M-token menu anyway? Four costs bite:
- Embedding-table cold start: each row is learned from that piece's occurrences. A token appearing 5 times in 15T training tokens gets a nearly-random vector. Huge vocabularies fill the long tail with undertrained entries that models then misuse or hallucinate IDs for — the "cold dishes" of the kitchen analogy.
- Softmax dilution and memory: logits are
Vfloats per position. At V = 250k, the logit tensor for one 128k-token sequence in fp32 runs to hundreds of GB without sharding — a real serving constraint in 2024–2026 long-context inference. - Diminishing returns: merge-based vocabularies follow Zipf logic. Past roughly 64k–128k for broad multilingual text, each additional merge covers vanishingly less corpus mass, while parameter cost keeps growing linearly.
- Tokenizer fairness: bigger vocabularies change tokenization statistics so aggressively that cross-model comparisons must switch to bytes-per-token (BPT) or bits-per-byte (BPB) rather than raw perplexity or tokens-per-word (see the Perplexity concept). Two models with different V are literally measuring "per position" over different amounts of text.
Rule of thumb to memorize:
Vocabulary cost is fixed and paid every forward pass; coverage benefit is variable and paid back in shorter sequences.
08.In Practice: Choosing V (2024-2026 Playbook)
How teams actually decide, in order:
- Target languages + corpus mass. English-only, budget-constrained encoder: 30k–64k. Production multilingual chat model: 128k–256k. Code-heavy assistants push the top of that range, because identifiers like
getOrNullAsynconly merge into single tokens when V is big enough to afford them. - Parameter budget check. V x d must stay a sane share of total capacity — many labs cap the vocab at ~10–15% of parameters. Tied embeddings (one matrix serving both doors) halve the cost: GPT-2 ties; LLaMA-3 unties.
- Measure coverage, not vibes. Compute tokens-per-character (or BPT) per script on held-out traffic. If Japanese costs 4x English on your candidate V, raise V or rebalance training-corpus sampling. Multilingual vocabularies allocate dedicated token mass to scripts — which is why Gemini can spend ~256k rows without hurting English.
- Algorithm pairing. BPE/WordPiece stop at the target size; Unigram (SentencePiece) starts oversized and prunes by likelihood, giving smoother coverage curves at large V — one reason T5, ALBERT, LLaMA-1/2, and Gemma picked it.
And back to the kitchen for the closing lesson: the menu is printed before the restaurant opens, bolted to the wall forever. LLaMA's jump from 32k (LLaMA-2) to 128,256 (LLaMA-3) was not a settings tweak — it re-baked the whole model. Choose V the way you would choose a building: cheap to plan, ruinous to change.
Architectural Trade-offs & Production Realities
Architectural Advantages
- Larger V: fewer tokens per text, shorter contexts, cheaper attention and better multilingual/code throughput.
- Whole-word coverage for common terms lets the model associate meaning with stable units.
- Bigger vocabularies reduce hallucinated-token artifacts from byte fallback on rare strings.
Trade-offs & Constraints
- Fixed cost: V x d parameters per edge matrix plus a V-wide softmax every position.
- Long-tail rows are starved of training signal, becoming dead or dangerous tokens.
- Cross-model metrics and costs become incomparable without bytes-per-token normalization.
LLaMA-3 expanded from LLaMA-2's 32k SentencePiece vocab to 128,256 byte-fallback BPE tokens, explicitly to cut token counts for code and multilingual text. OpenAI shipped o200k_base (~200k) with GPT-4o for the same reason. Google's Gemini line runs ~256k SentencePiece units with large script allocations. Each move bought real serving-cost reductions of 20–40% per language — a tokenizer change that behaves like a hardware upgrade.
Staff+ Engineering Takeaways
- Vocabulary size enters the model twice: the V x d embedding matrix and the V x d output projection / softmax.
- Small vocabularies fragment text into more tokens, inflating context use, latency, and per-language cost.
- Huge vocabularies pay fixed parameter and softmax cost and leave long-tail rows undertrained.
- The trend 2018-2024: 30k (BERT) -> 50k (GPT-2) -> 100k (GPT-4) -> 128k-256k (LLaMA-3, Qwen, GPT-4o, Gemini), driven by multilingual and code economics.
- Compare tokenizers with bytes-per-token/BPB, never raw token counts; V is frozen into a checkpoint and unchangeable after pre-training.
Topic Knowledge Check
Exercise 1 of 3 • Test your architectural comprehension.
Doubling the vocabulary size of a fixed Transformer model most directly doubles what?
How clear and actionable was this distributed systems breakdown?