NLP & Word Representations
Natural language processing as an engineering discipline:
All Topics in Phase 7
0 of 15 completedA neural network cannot read letters — it only eats integer IDs. A tokenizer chops text into small pieces and looks each piece up in the model's fixed vocabulary. This one first step decides what the model can learn, what it costs per request, and why it cannot count the r's in "strawberry".
BPE builds a tokenizer's vocabulary the way a group chat invents shorthand: watch the text, glue the most frequent pair of symbols into one new symbol, repeat until the vocabulary is full. Encoding new text is just replaying those glues in order. It is the algorithm behind GPT-family tokenizers.
WordPiece runs the same greedy merge loop as BPE but hires a stricter referee: a merge only earns a vocabulary slot if the pair appears meaningfully more often than its halves would predict by chance. It is the tokenizer behind BERT's famous 30,522-entry vocabulary.
Every tokenizer ships with one number nobody sees in the demo: the vocabulary size — 30k to 256k slots. That single choice sets embedding-table parameters, softmax cost, how many tokens each language burns per sentence, and how much a model costs to serve. Bigger is not free; smaller is not cheap either.
Words become points in a geometric space where "near" means "similar in meaning". An embedding is just a row of a learned matrix — and once meaning is numbers, similarity and even analogies like king - man + woman ≈ queen become arithmetic. This is the idea every modern embedding model descends from.
Word2Vec learns the meaning-map by playing a guessing game: predict a word from its neighbors (CBOW) or its neighbors from it (skip-gram). The vectors are a by-product of getting good at the game — and two sampling tricks turned an impossible softmax into a pipeline that trains on billions of words.
GloVe skips the guessing game and looks at the whole corpus at once: count how often every word appears near every other word, then train vectors so their dot products reproduce the log of those counts. "Counts, but also gradients" — global statistics meet learned geometry.
Static embeddings give every word one frozen address — "bank" the river and "bank" the loan live at the same coordinates. ELMo and then BERT fixed this by recomputing each token's vector from the whole sentence, trading a lookup for a forward pass and launching the entire modern NLP stack.
BERT is the 2018 paper that made "pre-train once, fine-tune everywhere" the default way to build NLP models. It is a Transformer encoder stack trained by fill-in-the-blank practice on BooksCorpus plus Wikipedia, released as 110M and 340M parameter checkpoints — and it swept every GLUE and SQuAD leaderboard.
Masked language modeling is the fill-in-the-blank trick that lets a model train on unlabeled text in both directions at once: corrupt about 15% of tokens with an 80/10/10 scheme, predict only the holes, and learn context from every side. With span variants like T5's sentinel tokens, it shapes how every modern encoder learns.
Next sentence prediction was BERT's second training task: show the model two sentences and ask, "does B really follow A?" It was bolted onto masked language modeling to give BERT document-level coherence — and it became the cautionary tale of the field, when RoBERTa's experiments showed the task was doing more harm than good.
The GPT lineage is one recipe — predict the next word — pulled through six upgrades: a decoder pre-trained then fine-tuned (GPT-1, 117M), zero-shot at 1.5B scale (GPT-2), few-shot prompting at 175B (GPT-3), human alignment (InstructGPT), multimodality (GPT-4/4o), and test-time reasoning (o1-o3). This topic walks how each step changed one lever and turned a language model into the default foundation for chat assistants, agents, and code generation.
Every LLM rests on one idea: write the probability of a whole sequence as a chain of next-token steps, P(x) = Π P(xₜ | x<ₜ). Train all the steps at once with a triangular attention mask and teacher forcing; run them one at a time at inference with a KV cache; then decide, by your sampling rule, how bold or careful the model sounds.
One attention block, three wiring choices: let every token see everything (encoder, a champion reader), let each token see only the past (decoder, a champion writer), or read the input fully and write the output while glancing back (encoder-decoder, a champion translator). The masking policy alone decides which tasks each family can win — understanding, generation, or transformation.
Perplexity turns a language model's average surprise into one human-scale number: exponentiate the mean negative log-likelihood and you get "roughly how many equally-plausible options the model thinks it faces at each step." Perfect model = 1, uniform guesser over K words = K. This topic covers the math, the entropy floor, the measurement traps, and exactly what the number can and cannot tell you.