Tokenization: From Raw Text to Model Input
A neural network cannot read letters — it only eats integer IDs. A tokenizer chops text into small pieces and looks each piece up in the model's fixed vocabulary. This one first step decides what the model can learn, what it costs per request, and why it cannot count the r's in "strawberry".
The Tokenization Pipeline 🔤
Text flows through progressively finer splitting stages. Word-level vocabularies fail on unknown words, which is why modern systems fall back to subword decomposition before embedding lookup.
01.The Problem: Computers Cannot Read Letters
You type a sentence into a chatbot:
"I love pizza"
But a neural network does not read strings. It consumes sequences of integer IDs. Numbers only.
So before any learning or predicting can happen, something must turn
text → numbers
That "something" is the tokenizer.
Here is the whole idea in one bold line:
A tokenizer chops raw text into small pieces, then looks each piece up in the model's fixed vocabulary — like pressing numbered slots on a vending machine.
A vending machine does not understand "cola". It understands slot 42.
A tokenizer is the person who reads your order and presses the buttons. The model's embedding table is the row of drinks behind those slots: each piece maps to a row index, and each row holds the numbers the network actually works with.
Getting text into that table is a deterministic, model-specific bridge between raw Unicode and the numeric representation the model consumes.
And this first step quietly controls everything downstream:
- Sequence length: how many tokens a sentence takes up decides context-window usage and inference cost. More tokens = slower and pricier.
- What the model can learn: the model only ever sees the units the tokenizer produces. The granularity literally shapes its notion of "word".
- Reproducibility: the same string must always encode to the same IDs. If the training tokenizer and the serving tokenizer disagree, quality degrades silently — no error, just a worse model.
One critical reality of 2024–2026: tokenizers are not interchangeable.
Sending text through the wrong tokenizer (a WordPiece encoder feeding a GPT-style model, say) produces garbage IDs with no error signal. The model still outputs fluent-looking text — just wrong.
So: everything a language model is, starts with how text becomes integers.
02.The Idea in Plain Words: Text Becomes Numbered Tickets
Strip away the jargon and a tokenizer does exactly two jobs:
Job 1: segment. Job 2: look up.
Segment — cut the raw character stream into discrete units ("tokens").
Look up — replace each unit with its slot number in a fixed vocabulary the tokenizer shipped with.
Let's unpack the pieces:
- A unit (token) can be a word, a piece of a word, a character, or a byte — depending on the tokenizer's strategy.
- The vocabulary is the fixed list of units this tokenizer knows. Typically 30k–256k entries for modern models.
- An ID is just the row number of the unit in that list.
- Decoding is the movie played backwards: IDs → units → glued back into text.
The trickier part: the vocabulary is decided before training and frozen forever after. The model can never learn a unit that is not on the menu.
That single constraint — a fixed menu — is the villain of this whole topic. Hold it in mind; we will meet it again as the "out-of-vocabulary problem".
03.A Tiny Worked Example: Encoding "The dog."
Imagine a toy tokenizer with a tiny vocabulary. Three of its slots happen to hold:
1996 → "The"
3814 → "dog"
6124 → "."
You feed it the sentence:
"The dog."
Step 1 — split on whitespace:
["The", "dog."]
Already we see a problem: dog. — word plus period glued together — is not in the vocabulary.
Step 2 — split on punctuation:
["The", "dog", "."]
Step 3 — look each piece up:
The → 1996
dog → 3814
. → 6124
The model receives the integer sequence:
[1996, 3814, 6124]
Done. Three characters of meaning, three button presses.
Now try a stranger word:
"unhappiness"
It is not in the vocabulary at all. A word-level tokenizer would collapse and hand back a generic <UNK> (unknown) token — all meaning gone.
A subword tokenizer instead chops it into known pieces:
["un", "happi", "ness"] → [4821, 17409, 9034]
One English word, three tickets.
Notice what just happened: for the first time, any string can be encoded — because you can always split finer. Remember the number 3 for "happiness": that price is exactly what drives API bills later.
04.Visual Intuition: The Pipeline
Picture text flowing left to right through progressively finer sieves:
coderaw text "The unhappiness dog ran!" ↓ [split spaces] [The] [unhappiness] [dog] [ran!] ↓ [split punct] [The] [unhappiness] [dog] [ran] [!] ↓ [in vocab?] ✓ ✗ OOV! ✓ ↓ ↓ [subword] [un][happi][ness] ↓ [look up] 1996 4821 17409 9034 3814 7642 2920 ↓ [embedding rows] ▓ ▓ ▓ ▓ ▓ ▓ ▓ ← one vector per ID
Each box only does one small thing.
And look at the shape of the funnel: the deeper the sieve falls (words → subwords → characters → bytes), the never-fails property improves — but the number of tickets grows.
That tension — coverage vs. compactness — is the entire design space of tokenization. Every strategy we meet next sits at a different point on this funnel.
05.Four Classical Ways to Load the Vending Machine
Before subwords took over, four families of tokenizers were the standard menu designs:
- Whitespace splitting: the naive
text.split(" "). Fast and dumb. It conflates"dog","dog.", and"(dog)"into different (or wrongly identical) strings — punctuation becomes noise baked into the tickets. - Rule-based / punctuation-aware tokenizers: the Penn Treebank and Berkeley Regular Expression tokenizers peel punctuation off as separate tokens. Moses, built for machine translation, added language-specific rules. To this day, Elasticsearch's "standard analyzer" (Unicode Text Segmentation, UAX #29) uses this style in production search stacks.
- Character-level tokenization: every glyph gets its own slot. Completely out-of-vocabulary-proof — the machine stocks only letter tiles, so any drink can be spelled. Great for noisy text and languages with complex word-building. The price: sequences grow 4–8x longer (imagine pressing c-o-l-a-and-enter for every order), and the model must learn word identity from scratch. Think char-LSTMs, byT5.
- Word-level tokenization: one slot per dictionary word. Compact and linguistically satisfying — the machine stocks every drink in the world, pre-mixed. But fatally brittle. See the next section.
Back to the vending machine: strategy 3 is a machine that only sells LEGO bricks; strategy 4 is a machine that must stock a slot for every finished product ever invented. You can already feel why industry eventually picked the middle.
One more wrinkle before we move on: Chinese and Japanese have no spaces at all.
我爱披萨
Whitespace splitting finds exactly one "word": the whole sentence. So for these languages a word-segmentation step must run before any word-level model — jieba for Chinese, MeCab for Japanese. And if the segmenter slices in the wrong place, the error is baked into every downstream ID, unrecoverable.
06.The Villain: Out-of-Vocabulary Disasters
Strategy 4 looks lovely until real text shows up.
A fixed word-level vocabulary can never cover real-world text. New products, slang, typos, transliterations, and code identifiers appear constantly.
An unknown word — say "quarknetbook" — gets no embedding. Typically it collapses into a generic <UNK> token, which destroys the meaning of the sentence:
"I bought a quarknetbook" → "I bought a <UNK>"
The model learned nothing about the one word that mattered.
And the problem compounds in three directions:
- Morphologically rich languages: Turkish, Finnish, and Russian generate millions of word forms from small roots. Turkish
"evlerinizden mi"literally packs "is it from your houses?" into a few fused pieces. Any feasible dictionary misses legitimate forms a native speaker uses daily. - Agglutinative and compounding languages: German glues nouns into monsters —
"Rindfleischetikettierungsüberwachungsaufgabenübertragungsgesetz"(yes, a real historic law name). One "word", zero dictionary slots. - Non-language content: numbers like
"1,234,567", URLs, emojis, and code are strings a word-based tokenizer was never meant to be fluent in.
Worse, word-level vocabularies force models to treat "dog", "dogs", and "doggedly" as totally unrelated symbols — the shared "dog" meaning is invisible to the machine, wasting parameters and blocking knowledge transfer between obviously related forms.
So the question becomes:
Can the vending machine ever stock enough slots?
Not with words. But there is a way to make the menu small and infinite at the same time...
07.The Subword Compromise — and Its Weird Side Effects
Since about 2016, every serious NLP system — BERT, GPT, T5, LLaMA, Gemini, Qwen, DeepSeek — converged on the same answer: subword tokenization.
One fixed vocabulary of 30k–256k pieces — mixing whole words, common morphemes, and characters — learned from corpus statistics.
Instead of asking a human to write the menu, the machine watches a giant pile of text and promotes the pieces that keep reappearing.
The three dominant algorithms for learning that menu — Byte-Pair Encoding (BPE), WordPiece, and Unigram/SentencePiece — each get their own concept next, including a hands-on merge-by-merge example.
Subwords kill the OOV problem: any string, no matter how weird, decomposes into known pieces. And statistics are shared across related words: dog, dogs, and doggedly now literally share a token.
But the compromise introduced odd artifacts that LLM engineers live with in 2024–2026:
- Tokens ≠ words: English averages roughly 1.3 tokens per word. German or Japanese text can cost 3–7x more tokens for the same amount of meaning — which changes API billing and multilingual capability directly. (That worked example "un-happi-ness"? Three tickets, one word.)
- Number and spelling blind spots:
"9.11"vs"9.9"tokenize differently across versions, so models compare them badly; and models famously fail to count the letters in "strawberry" — they never see individual letters as reliable units, only fragments. - Tokenizer-level bugs: prompt injection and repetition loops can be triggered by boundary effects and byte-roundtrip failures in byte-level vocabularies. Some "AI misbehavior" is really a button-pressing bug.
Same model, different machine:
code"Tokenization is half the battle in NLP!" BERT menu: token | ##ization | is | half | the | battle | in | nl | ##p | ! GPT-2 menu: Token | ization | Ġis | Ġhalf | Ġthe | Ġbattle | Ġin | ĠN | LP | !
Different vocabularies slice the same sentence into different numbers of tickets — and each model only understands its own menu.
from transformers import AutoTokenizer
tok_bert = AutoTokenizer.from_pretrained("bert-base-uncased") # WordPiece, 30k vocab
tok_gpt = AutoTokenizer.from_pretrained("gpt2") # Byte-level BPE, 50k vocab
text = "Tokenization is half the battle in NLP!"
print(tok_bert.tokenize(text)) # ['token', '##ization', 'is', 'half', 'the', 'battle', 'in', 'nl', '##p', '!']
print(tok_gpt.tokenize(text)) # ['Token', 'ization', 'Ġis', 'Ġhalf', 'Ġthe', 'Ġbattle', 'Ġin', 'ĠN', 'LP', '!']
# Note: BERT lower-cases and uses ## markers; GPT-2 preserves case and hides spaces inside tokens ('Ġ').08.In Practice: Live With the Menu You Trained On
Three rules engineers learn the hard way:
1. The tokenizer is welded to the checkpoint. Embedding rows and token IDs are learned together. ID 4213 means one piece in BERT's vocabulary and something unrelated in GPT's. Serving with the wrong tokenizer produces silent garbage — always ship the exact tokenizer paired with the model (OpenAI's tiktoken exists precisely for this).
2. Count tokens before you count parameters. Your bill, your context window, and your latency all tick in tokens, not words. A multilingual product on an English-heavy vocabulary can pay 4–6x per request for the same meaning — the tokenizer choice is a pricing decision.
3. Expect weirdness at the seams. Letter counting, spelling puzzles, number comparison, exact-case reconstruction: these "dumb model failures" are often tokenizer artifacts from the funnel in section 4 — the information was already thrown away at button-pressing time.
The vending machine framing carries through the whole phase:
BPE, WordPiece, and Unigram are just three different ways of deciding which drinks get their own slot — and vocabulary size is the rent you pay for the machine.
Next: exactly how BPE learns its menu by merging.
Architectural Trade-offs & Production Realities
Architectural Advantages
- Whitespace and rule-based tokenizers are fast, transparent, and linguistically interpretable.
- Character-level tokenization is 100% OOV-free and robust to typos and noise.
- Subword tokenization balances coverage, compactness, and statistical sharing across word families.
Trade-offs & Constraints
- Word-level vocabularies explode on morphologically rich languages and never achieve full coverage.
- Character-level sequences are long, slow to train, and weaken attention resolution.
- Subword tokenizers fragment rare words, distort number comparison, and complicate exact string tasks.
Elasticsearch ships per-language analyzers (standard Unicode segmentation + stemming + stop words) so inverted-index term matching stays deterministic and debuggable. OpenAI pairs each model with an exact byte-level BPE tokenizer in tiktoken (cl100k_base for GPT-3.5/4, o200k_base for GPT-4o) because decoding the wrong way changes billing, truncation behavior, and model quality.
Staff+ Engineering Takeaways
- Models consume integer IDs; the tokenizer is the model-specific, irreversible bridge from Unicode text to those IDs.
- Word-level tokenization fails on out-of-vocabulary words and morphologically rich languages.
- Character-level tokenization eliminates OOV but multiplies sequence length and learning difficulty.
- Subword algorithms (BPE, WordPiece, Unigram) are the universal modern compromise, learned from corpus statistics.
- Tokens are not words: tokenization granularity directly shapes cost, multilingual coverage, and LLM failure modes.
Topic Knowledge Check
Exercise 1 of 3 • Test your architectural comprehension.
Character-level tokenization never hits an unknown word. Why is it still almost never the main strategy in modern LLMs?
How clear and actionable was this distributed systems breakdown?