TOPIC #129Intermediate 15 min read

Vocabulary Size: Coverage vs. Compute in Tokenizer Design

AI
AI & ML Editorial
Report an issue
Key takeawayCore Concept Summary

Every tokenizer ships with one number nobody sees in the demo: the vocabulary size — 30k to 256k slots. That single choice sets embedding-table parameters, softmax cost, how many tokens each language burns per sentence, and how much a model costs to serve. Bigger is not free; smaller is not cheap either.

Vocabulary Size Is a Two-Sided Lever ⚖️

Growing the vocabulary buys coverage and shorter sequences, but charges you in fixed parameters, softmax compute, and diluted training signal for rare pieces.

Vocabulary Size Is a Two-Sided Lever ⚖️
100%
Touchpad: Pinch to zoom • Drag to pan
Rendering visual architecture flowchart...

01.The Problem: One Number Nobody Puts on the Box

You have now met the vending machine (tokenization) and its menu-writing algorithm (BPE, WordPiece, Unigram).

Every menu has a length. The tokenizer for BERT carries 30,522 pieces. GPT-4o carries roughly 200,000. Gemini carries about 256,000.

How was that number chosen? And why should you care?

Insight

Because the vocabulary size — usually written V — silently controls four things at once:

  1. how many parameters the model burns on the first and last layer,
  2. how expensive every forward pass is,
  3. how many tokens (money!) each language costs per sentence,
  4. how well rare pieces are actually learned.

Pick V too small and everything gets chopped into too many pieces — Japanese text can cost 3–6x more tokens than it should, and every one of them is billed.

Pick V too big and you pay a fixed tax in weights and compute on every single step — plus a long tail of vocabulary rows that barely trained at all.

This concept is about that trade. It is a two-sided lever, and once trained, the lever is welded in place.

02.The Idea in Plain Words: V Is the Menu Length

First, where does V actually live inside a Transformer? Almost nowhere — exactly two places, both at the edges of the network:

  1. Input embedding matrix: shape V x d. V rows, each a d-dimensional vector. (From the embeddings concept: a token's ID just picks its row.)
  2. Output projection ("unembedding"): another V x d matrix, unless tied with the input. Its job is to score every vocabulary item at every step so the softmax — the layer that turns scores into probabilities — can pick the next token.

Everything in the middle (attention heads, MLPs, layers) does not care about V at all.

So the mental model:

Insight

V sets the size of the model's front door and back door. The hallway in between is built to a different spec.

Front door: numbers → vectors. Back door: vectors → probabilities over all V candidates.

Let's put real numbers on the doors.

03.A Tiny Worked Example: Counting Doors

Part A — the parameter bill.

LLaMA-3 has V = 128,256 and hidden size d = 4,096.

V x d = 128,256 x 4,096 ≈ 525,000,000

That is ~0.53 billion parameters in the embedding matrix alone. LLaMA-3 unties its output projection, so the back door is another 0.53B.

Insight

0.53B + 0.53B ≈ over 1 billion parameters for a model whose total is 8B.

More than an eighth of the entire model is the menu.

Part B — the token bill.

Take the word internationalization.

  • Under a small 30k English-leaning vocabulary it might shatter into inter + na + tion + al + ization → 5 tickets.
  • Under a 200k multilingual vocabulary it is plausibly 1 ticket.

Same meaning. Five times the positions. And positions are where everything hurts: attention cost grows quadratically with sequence length, context windows fill up 5x faster, and per-token API pricing multiplies by five.

Same lever, opposite ends: Part A is the fixed cost of a big menu, Part B is the variable cost of a small one.

python— Counting vocabulary parameters in a shipped checkpoint
# LLaMA-3-8B: 128,256 vocab x 4096 hidden, untied embeddings
V, d = 128_256, 4096
print(f"Embedding table:  {V*d/1e9:.2f}B params")   # 0.53B
print(f"Unembedding:      {V*d/1e9:.2f}B params")   # 0.53B (untied)
# vs bert-base-uncased: 30,522 x 768 = 0.023B — tiny because encoders
# classify with small heads instead of predicting over the full vocab

04.Visual Intuition: A Seesaw, and a Timeline

The whole topic balances on one seesaw:

code
        SMALL V                          BIG V
   ┌────────────────┐              ┌────────────────┐
   │ tiny doors     │              │ 0.5B+ params   │ ← fixed tax,
   │ cheap softmax  │              │ fat logit layer│   every pass
   │                │              │                │
   │ every rare word│              │ common words   │
   │ shatters into  │  ◄──seesaw──►│ fit in ONE     │
   │ many tokens    │              │ token          │
   │ (3-6x for      │              │                │
   │  Japanese!)    │              │ long tail of   │
   └────────────────┘              │ cold rows      │
                                   └────────────────┘

And the industry has been visibly sliding right along that lever:

code
2018    BERT          30k
2019    GPT-2         50k
2023    GPT-3.5/4     100k
2024    GPT-4o        200k     LLaMA-3  128k
2024+   Qwen          152k     Gemini  ~256k
        ────────────────────────────────────────►
        vocab sizes grew ~5–10x in six years,
        driven by multilingual + code economics

Nobody did this out of greed for big numbers. Each jump paid for itself in shorter sequences.

05.The Analogy: A Kitchen That Pre-Cooks Everything

Carry one picture through the rest of this concept: a restaurant kitchen deciding how many dishes to pre-cook.

The vocabulary is the set of pre-made dishes. A sentence is an order the kitchen assembles by platter-combining.

  • Many pre-made dishes (big V): most orders need only two or three platter grabs — service is fast, the counter (context window) empties quickly. But the rent is brutal: every dish needs its own cabinet slot (an embedding row), and the waiter must check every dish before recommending one (softmax over V).
  • Few pre-made dishes (small V): rent is cheap. But every order of "internationalization" requires scrambling five ingredients together from basic stock — slower service, longer lines, and the same popular dish arrives at the table in fragments.
  • Rare dishes (the long tail): a dish ordered 5 times all season never gets practiced. Its cook (the embedding row) learns almost nothing from those 5 orders — it stays close to raw ingredients, i.e., a near-random vector the model then misuses.

Three bills, one kitchen. Keep them separate:

Insight

Rent = fixed parameters + softmax width. Prep time = tokens per text. Practice = occurrences per vocabulary row.

Growing V cuts prep time, raises rent, and dilutes practice.

06.Why AI Cares — The Coverage Side of the Ledger

A vocabulary too small for its training distribution fragments text: common words get split into several pieces, so every sentence costs more tokens. The measured numbers:

  • English under a 30k WordPiece vocabulary averages ~1.4–1.6 tokens per word; under a modern 128k–200k byte-level BPE, ~1.2–1.3.
  • For morphologically rich or non-Latin languages the gap explodes: Japanese or Chinese under an English-heavy 50k vocabulary can run 3–6x more tokens than under a 128k+ multilingual vocabulary.
  • Longer sequences mean: quadratic attention cost, faster context-window exhaustion, slower inference, and higher price per unit of meaning. This is why tokenizer choice was the headline complaint about early English-centric models serving global traffic.

The trend line in section 4 is this ledger in motion: BERT (2018) 30k → GPT-2 (2019) 50k → GPT-3.5/4 (2023) 100k → GPT-4o (2024) 200k, LLaMA-3 (2024) 128k, Qwen 152k, Gemini ~256k — vocabularies grew ~5–10x in six years, driven almost entirely by multilingual and code economics.

07.Why AI Cares — The Cost Side of the Ledger

Compute per step also scales with V. The final projection produces a V-wide logit vector and the softmax cross-entropy sums over all of it. For large V and long contexts, frameworks shard this computation across GPUs (tensor/vocabulary parallelism, e.g., Megatron-LM) precisely because the V x d matmul is the single widest linear layer in the stack.

Why not just ship a 1M-token menu anyway? Four costs bite:

  • Embedding-table cold start: each row is learned from that piece's occurrences. A token appearing 5 times in 15T training tokens gets a nearly-random vector. Huge vocabularies fill the long tail with undertrained entries that models then misuse or hallucinate IDs for — the "cold dishes" of the kitchen analogy.
  • Softmax dilution and memory: logits are V floats per position. At V = 250k, the logit tensor for one 128k-token sequence in fp32 runs to hundreds of GB without sharding — a real serving constraint in 2024–2026 long-context inference.
  • Diminishing returns: merge-based vocabularies follow Zipf logic. Past roughly 64k–128k for broad multilingual text, each additional merge covers vanishingly less corpus mass, while parameter cost keeps growing linearly.
  • Tokenizer fairness: bigger vocabularies change tokenization statistics so aggressively that cross-model comparisons must switch to bytes-per-token (BPT) or bits-per-byte (BPB) rather than raw perplexity or tokens-per-word (see the Perplexity concept). Two models with different V are literally measuring "per position" over different amounts of text.

Rule of thumb to memorize:

Insight

Vocabulary cost is fixed and paid every forward pass; coverage benefit is variable and paid back in shorter sequences.

08.In Practice: Choosing V (2024-2026 Playbook)

How teams actually decide, in order:

  1. Target languages + corpus mass. English-only, budget-constrained encoder: 30k–64k. Production multilingual chat model: 128k–256k. Code-heavy assistants push the top of that range, because identifiers like getOrNullAsync only merge into single tokens when V is big enough to afford them.
  2. Parameter budget check. V x d must stay a sane share of total capacity — many labs cap the vocab at ~10–15% of parameters. Tied embeddings (one matrix serving both doors) halve the cost: GPT-2 ties; LLaMA-3 unties.
  3. Measure coverage, not vibes. Compute tokens-per-character (or BPT) per script on held-out traffic. If Japanese costs 4x English on your candidate V, raise V or rebalance training-corpus sampling. Multilingual vocabularies allocate dedicated token mass to scripts — which is why Gemini can spend ~256k rows without hurting English.
  4. Algorithm pairing. BPE/WordPiece stop at the target size; Unigram (SentencePiece) starts oversized and prunes by likelihood, giving smoother coverage curves at large V — one reason T5, ALBERT, LLaMA-1/2, and Gemma picked it.

And back to the kitchen for the closing lesson: the menu is printed before the restaurant opens, bolted to the wall forever. LLaMA's jump from 32k (LLaMA-2) to 128,256 (LLaMA-3) was not a settings tweak — it re-baked the whole model. Choose V the way you would choose a building: cheap to plan, ruinous to change.

Architectural Trade-offs & Production Realities

Architectural Advantages

  • Larger V: fewer tokens per text, shorter contexts, cheaper attention and better multilingual/code throughput.
  • Whole-word coverage for common terms lets the model associate meaning with stable units.
  • Bigger vocabularies reduce hallucinated-token artifacts from byte fallback on rare strings.

Trade-offs & Constraints

  • Fixed cost: V x d parameters per edge matrix plus a V-wide softmax every position.
  • Long-tail rows are starved of training signal, becoming dead or dangerous tokens.
  • Cross-model metrics and costs become incomparable without bytes-per-token normalization.
Production Implementation in Big Tech
Meta / OpenAI / Google• The 2023-2024 vocabulary land-grab

LLaMA-3 expanded from LLaMA-2's 32k SentencePiece vocab to 128,256 byte-fallback BPE tokens, explicitly to cut token counts for code and multilingual text. OpenAI shipped o200k_base (~200k) with GPT-4o for the same reason. Google's Gemini line runs ~256k SentencePiece units with large script allocations. Each move bought real serving-cost reductions of 20–40% per language — a tokenizer change that behaves like a hardware upgrade.

Staff+ Engineering Takeaways

  • Vocabulary size enters the model twice: the V x d embedding matrix and the V x d output projection / softmax.
  • Small vocabularies fragment text into more tokens, inflating context use, latency, and per-language cost.
  • Huge vocabularies pay fixed parameter and softmax cost and leave long-tail rows undertrained.
  • The trend 2018-2024: 30k (BERT) -> 50k (GPT-2) -> 100k (GPT-4) -> 128k-256k (LLaMA-3, Qwen, GPT-4o, Gemini), driven by multilingual and code economics.
  • Compare tokenizers with bytes-per-token/BPB, never raw token counts; V is frozen into a checkpoint and unchangeable after pre-training.

Topic Knowledge Check

Exercise 1 of 3 • Test your architectural comprehension.

Exercise 1 of 30 answered
1

Doubling the vocabulary size of a fixed Transformer model most directly doubles what?

Rate This Architecture ChapterFeedback & Rating

How clear and actionable was this distributed systems breakdown?