Word Embeddings: Meaning as Geometry
Words become points in a geometric space where "near" means "similar in meaning". An embedding is just a row of a learned matrix — and once meaning is numbers, similarity and even analogies like king - man + woman ≈ queen become arithmetic. This is the idea every modern embedding model descends from.
From Token IDs to Dense Vectors 🧭
An embedding is just a row lookup in a learned V x d matrix. Once words are points in a shared space, similarity and analogy become arithmetic.
01.The Problem: One-Hot Vectors Are a Dead End
The tokenizer did its job (previous concepts). You have integer IDs.
Now the ID has to enter a neural network, and networks eat vectors — lists of numbers. The textbook-first idea is the one-hot vector: with a 50,000-word vocabulary, "cat" is a vector of 50,000 zeros with a single 1 at position, say, 7,312:
codecat → [0, 0, ..., 1, 0, ..., 0] ← the 1 sits at position 7,312 dog → [0, 0, ..., 1, 0, ..., 0] ← the 1 sits at position 22,501
Two fatal problems:
- Dimensionality: every layer input is 50k-wide regardless of sentence content. Parameters explode, everything is sparse, and almost all of that width is zeros carrying nothing.
- No similarity structure: the dot product between any two distinct one-hots is exactly 0. cat · dog = 0. cat · "quarterly earnings" = 0. To the network, a cat is exactly as related to a dog as to a spreadsheet.
"cat" and "dog" being maximally unrelated is not a detail. It is the whole game.
Language runs on similarity — synonyms, categories, related topics. One-hot encodes the identity of a word and throws away every relationship.
So the question becomes:
Can we give every word a vector where distance means something?
02.The Idea in Plain Words: You Shall Know a Word by Its Company
In 1957, linguist J. R. Firth offered the seed of the answer:
"You shall know a word by the company it keeps."
This is the distributional hypothesis: words used in similar contexts have similar meanings.
- You meet "puppy" in sentences about fetching, barking, vet.
- You meet "dog" in sentences about fetching, barking, vet.
- Therefore: probably cousins.
Word embeddings operationalize the hypothesis:
Each word gets a dense vector of perhaps 50–300 real numbers, learned so that contextually similar words land near each other in space.
"Dense" just means all the entries are actual numbers, not zeros. A small list of coordinates, like a latitude/longitude pair — except the map has 300 dimensions and the axes were never named by a human.
And that reframes everything: semantics becomes geometry.
- Distance is relatedness.
- Direction can be a meaning relation.
- And geometry is differentiable — so meaning becomes trainable, and transferable between tasks.
That last bullet deserves one plain sentence, because it is the quiet revolution: with one-hot vectors, "learning about cats" teaches the network nothing about dogs (the 1 sits in an unrelated slot). With vectors, nudging the cat coordinate moves it toward or away from dog, kitten, fur — one gradient step ripples through every similarity relation at once. Statistics get shared, and that is why small dense vectors beat huge sparse ones.
The map metaphor is not decorative. Carry it through: from here on, every word has an address in meaning-space.
03.What an Embedding Actually Is (It Is Only a Row)
Mechanically, a word embedding layer is one matrix: E ∈ R^(V x d) — V rows, one per vocabulary entry, each row a d-dimensional vector.
"Embedding token 4213" means: slice out row 4213.
That is literally all. Equivalently, it is the one-hot vector multiplied by E — the sparse 50k-wide ticket gets compressed into a dense learned lookup.
So an embedding layer computes nothing clever at runtime. It is a constant-time row fetch. The entire intelligence lives in the values, which were learned from corpus statistics.
What does the learned map look like? Typical geometric properties after training:
- Cosine similarity cos(u, v) = (u · v) / (‖u‖ ‖v‖) in [-1, 1] measures topical/semantic relatedness: about 0.8 for "ocean/sea", around 0.1 for "ocean/mortgage". (Cosine = the angle between two arrows; same direction → 1, unrelated → 0.)
- Nearest neighbors of "hospital" come back as "clinic", "ward", "emergency" — things-often-found-with (syntagmatic) and things-of-the-same-kind (paradigmatic) relations mix together in one space.
- Rough axes emerge for number, tense, gender, and sentiment — no axis was supervised; they fall out of the prediction objective.
The dimension d is a knob: 300 was the word2vec/GloVe era standard (2013–2018); modern first-layer embeddings for LLMs run 1024–16k (and remember from the vocabulary-size concept: those rows are exactly the V x d embedding matrix, one per token).
import numpy as np
from gensim.test.utils import datapath
from gensim.models import KeyedVectors
kv = KeyedVectors.load_word2vec_format(datapath("glove_king_100d.csv"), binary=False, unicode_separator=False)
v = kv["king"] # row of E for token "king" -> np.ndarray shape (100,)
print([w for w, _ in kv.similar_by_word("ocean", topn=4)]) # sea, water, coastal, waters
# Analogy arithmetic:
best, _ = kv.most_similar(positive=["king", "woman"], negative=["man"], topn=1)
print(best) # 'queen' (cosine ranking over the whole space)04.Tiny Worked Examples: Two-Dimensional Meaning-Space
Shrink the map to 2 axes so you can see it: x = maleness, y = royalty (each from 0 to 1). Suppose training produced:
codeking = (0.9, 0.9) woman = (0.2, 0.5) man = (0.9, 0.2) queen = (0.2, 0.8) cat = (0.3, 0.1)
Analogy arithmetic. Walk the famous equation:
vec(king) − vec(man) + vec(woman)
= (0.9−0.9+0.2, 0.9−0.2+0.5)
= (0.2, 1.2)
Which real word is closest? Distances: to queen (0.2, 0.8) → 0.4. To woman (0.2, 0.5) → 0.7. To king → 0.76.
queen wins. Subtract "maleness", add "femaleness", land near the queen.
(Real embeddings do this with 300 axes and hundreds of tangled relations — but the arithmetic is the identical vector add/subtract, then a nearest-neighbor search.)
Cosine similarity. Even smaller: with cat = (1, 2):
dog = (2, 1): cos = (1·2 + 2·1) / (√5 · √5) = 4/5 = 0.8 → nearly same direction, similar company.car = (2, −1): cos = (1·2 + 2·(−1)) / (√5 · √5) = 0/5 = 0.0 → perpendicular, unrelated.
Compare with one-hot: cat · dog = 0 and cat · car = 0 — both "maximally different". The dense map immediately does the thing the one-hot dead end could not.
05.Visual Intuition: The Neighborhood Map
Draw the meaning-city. Positions are learned, axes are unnamed — but the layout tells the story:
coderoyalty ▲ │ king ● ● queen │ │ │ man ● ● woman │ │ ● president │ └──────────────────────────► (maleness →) ● king − man + woman: start at king, step "toward commoner", then "toward female" → lands on queen ocean ●───≈───● sea ● mortgage (short walk = similar meaning) (far = unrelated)
Two readings of the same picture:
- Distance = relatedness. Neighbors share meaning; the suburbs of "bank" and "mortgage" are close, the downtown of "bank" and "sandwich" is a long commute.
- Consistent arrows = relations. The step from man to woman, king to queen, bachelor to spinster — the same arrow shape repeated across the map. That is why vector arithmetic can work.
06.The Famous Vector Algebra — and Its Real Limits
The headline result that made embeddings famous, from Mikolov et al. (2013): with trained embeddings,
vec("king") − vec("man") + vec("woman") ≈ vec("queen")
and analogous regularities held for capital-cities (Athens − Greece + Italy ≈ Rome), comparatives, and verb tenses.
Plausible explanation, in map terms: consistent semantic relations correspond to roughly translation directions — parallel arrows scattered across the city — so relation arithmetic works when the relation axis is approximately linear.
Important 2024–2026 nuance, though. Later analyses (Dziri et al., 2016, "Is this the observed reality?") showed the original analogy benchmarks were partly artifacts of training objectives and evaluation quirks — analogy accuracy is much weaker on clean, lemmatized corpora.
The honest summary:
The phenomenon is real but bounded. Vector algebra captures some relations, not "meaning as linear algebra". Treat analogy demos as pedagogical intuition, not engineering guidance.
07.In Practice: How the Coordinates Get Learned (and the Static Ceiling)
Nobody hand-marks "royalty" values and supervises the map. The coordinates are learned as a by-product of a prediction task:
- Skipgram/CBOW (Word2Vec): predict a word's neighbors (or vice versa) inside a context window; the hidden-layer weights become the vectors. (Next concept, with the full architecture.)
- GloVe: regress the whole word-word co-occurrence matrix with a factorization objective. (Concept after next.)
- Static features era (2014–2018): concatenate or average word vectors into LSTMs/CNNs for sentiment and classification — the "just add the numbers" phase of industrial NLP.
(And recall from the tokenizer concepts: these tables live inside a V x d matrix, so vocabulary size directly sets how many addresses the city contains.)
The structural limitation of everything above, in one line:
One vector per word form, forever.
Consequences that will not stay buried:
"bank"in "river bank" and "bank account" gets the identical vector — one frozen address for two different places.- Word order is invisible: "dog bites man" and "man bites dog" are the same bag of vectors.
- Negation is invisible: "not good" sits near "good".
- Polysemy shows up as clustered/dominated vectors: static embeddings average senses, producing geometric mush for ambiguous words — a measurable pathology, not just a philosophical one.
Fixing exactly this motivated contextualized embeddings — ELMo and then BERT (Concept 133) — where the "embedding" of a token is recomputed fresh from its entire sentence.
The map survives; the map just becomes a live traffic map instead of a printed one. That transition is the next several concepts — and the reason everything you just read is the foundation of RAG, search, and recommendation stacks today.
Architectural Trade-offs & Production Realities
Architectural Advantages
- Dense vectors generalize: similar contexts yield similar numbers, enabling transfer across tasks.
- Lookup is O(1); embeddings are tiny compared to any downstream network.
- Pre-trained vectors (word2vec/GloVe) work out of the box for similarity, clustering, and features.
Trade-offs & Constraints
- One static vector per word form cannot separate senses (polysemy) or model word order.
- Analogy directions are approximate and dataset-sensitive, not a reliable semantic algebra.
- Averaging-based sentence features discard structure; contextual models are needed for SOTA quality.
Before LLMs, production systems embedded entities the same way: playlists and airports as "words" in pseudo-sentences of co-listening/co-visit sequences, trained with word2vec-style objectives. Spotify published this exact recipe in 2017 — 100-dimensional static vectors, cosine nearest neighbors for radio and similarity — the direct ancestor of today's item-embedding stacks.
Staff+ Engineering Takeaways
- Word embeddings replace sparse, similarity-blind one-hot vectors with dense learned vectors grounded in the distributional hypothesis.
- Mechanically, embedding = slicing row i of a learned V x d matrix; the meaning lives in values learned from a prediction objective.
- Cosine geometry makes similarity, clustering, and relation arithmetic (king-queen) possible in vector space.
- Analogy linearity is real but bounded and partly benchmark-dependent.
- The static one-vector-per-word limit (polysemy, no word order) is the bridge to Word2Vec/GloVe mechanics and, decisively, contextual embeddings.
Topic Knowledge Check
Exercise 1 of 3 • Test your architectural comprehension.
What does the embedding layer compute at runtime for token index i?
How clear and actionable was this distributed systems breakdown?