Vector Embeddings for Retrieval
A retrieval embedding gives every piece of text a set of coordinates in a learned "meaning map", so texts with similar meaning land close together. Search then becomes simple geometry: rank by angle. This topic covers how those maps are trained, measured, and chosen.
From Contrastive Training to Cosine Ranking
Embedding models learn a geometry where semantic relevance equals angular proximity; retrieval is then nearest-neighbor search in that space.
01.The Problem: Words Are Not Meanings
You build a help-center search box.
A user types:
"how do I get back into my account when my code expired"
The correct article exists. Its title is:
"Resetting multi-factor authentication after TOTP timeout"
An old keyword search compares letters. Almost nothing matches: "code" ≠ "TOTP", "get back into" ≠ "resetting". The search shrugs and shows nothing.
So the question becomes
How can a computer measure what two texts are about, not just which characters they share?
The trick is to stop treating text as strings and start treating it as positions in space.
If we could hand every sentence a set of coordinates — like a GPS address — and arrange it so that sentences with similar meanings get nearby addresses, then "finding related text" would just mean "find nearby points." Geometry does the thinking.
That coordinate assignment is called an embedding, and this topic is about how those coordinates are made, what makes them good, and how to pick one.
02.The Idea in Plain Words: Text Becomes Coordinates
A retrieval embedding is simply
A fixed-length list of numbers assigned to a piece of text, arranged so that similar meanings get similar numbers.
Formally: the model maps variable-length text to a vector of typically 384–3072 dimensions (each dimension is one number in the list; you cannot see the whole vector in your head, but arithmetic on it is trivial).
- "kid", "child", "toddler" → three vectors sitting close together.
- "kid", "goat" → the word "kid" also means young goat, so it sits somewhere between both neighborhoods. Meaning clusters, not words.
How do the coordinates get arranged "correctly"? Nobody hand-labels positions. The model learns them through contrastive learning (Sentence-BERT, 2019; the E5/BGE/GTE families since):
Pull positives together, push negatives apart — repeat a billion times.
Training shows triplets:
- an anchor: a query text,
- a positive: a passage that genuinely answers it,
- many negatives: passages that do not.
The objective, InfoNCE, raises the cosine similarity of (anchor, positive) and lowers it for (anchor, each negative). Modern models mine hard negatives — plausible-but-wrong passages dug up by BM25 (an older word-matching scorer, topic 165) or earlier model checkpoints — because easy negatives teach nothing. This is why 2024–2026 models leapfrog older ones on retrieval, not just classification.
Three practical properties to burn in:
- Similarity is angular. Cosine similarity measures the angle between two vectors; and once vectors have length 1 (L2-normalized), cosine and the dot product give the same ranking. Most stacks normalize once at write time and use the cheaper inner product at query time.
- Dimensions are a budget, not a destiny. A 1536-d and a 3072-d model can both be good; quality comes from training data and objective, not raw size.
- Embeddings are model-scoped. Vectors from two different models are mutually meaningless — like coordinates on two different maps. Changing model means re-embedding the entire corpus.
03.A Simple Worked Example: Ranking by Angle
Forget 1536 dimensions. Use 2-D arrows so you can see every step.
Suppose our (toy) embedder outputs unit-length vectors, meaning length 1:
codequery q = (1.00, 0.00) doc A d₁ = (0.98, 0.17) ← nearly same direction doc B d₂ = (0.17, 0.98) ← almost perpendicular
Cosine similarity:
cos(a, b) = (a·b) / (‖a‖·‖b‖)
where a·b = a₁b₁ + a₂b₂ (multiply matching positions, add up) and ‖a‖ is the vector's length.
Because every vector here has length 1, the denominator is just 1, so cosine = dot product:
cos(q, d₁) = 1.00·0.98 + 0.00·0.17 = 0.98— doc A is about the same thing.cos(q, d₂) = 1.00·0.17 + 0.00·0.98 = 0.17— doc B is unrelated.
Rank by that number, take the top-K, done. Retrieval became arithmetic.
Now the normalization footnote matters in real life. Suppose lengths are not cleaned up: doc A becomes (0.49, 0.09) and doc B becomes (3.4, 19.6) — same directions, random magnitudes. The raw dot products with q are 0.49 vs 3.4: the unrelated-but-long doc B now wins. Magnitude noise flipped the ranking.
Divide each by its length and the angle survives: cos(q, d₁) = 0.49/0.5 = 0.98, cos(q, d₂) = 3.4/19.9 ≈ 0.17 — correct order restored. Only the angle should count. That is exactly why production code normalizes vectors once at write time and then ranks by the cheaper inner product.
04.Visual Intuition: The City of Meanings
Picture embedding space as a map — every text is an address on it.
code↑ dimension 2 │ "TOTP │ "password" reset"● ● │ ● q ← your query │ ● "2FA codes" ─────────┼──────────────→ dimension 1 │ │ ● "recipes for bread" ← far away, │ different neighborhood
- The query
qlands inside a cluster of authentication articles. Retrieval = "who lives nearest?" - Topic neighborhoods form naturally: support docs, recipes, legal clauses each occupy a district.
- Contrastive training is the city planner: it drags positives into the same block and shoves negatives to the opposite side of town.
The "unit sphere" version of the same idea: after normalization every vector touches the surface of one ball, and similarity is purely the angle at the center — the two pointers on a compass, agreeing or disagreeing.
05.The Analogy: GPS for Text — But Which Map?
Carry this analogy through the rest of the topic: embeddings are GPS coordinates for meaning.
- A good embedder makes addresses useful: nearby addresses really do mean nearby topics. Your phone can then say "find me something like this" by looking a few blocks away.
- A bad embedder is a mislabeled map: everything shifted, neighborhoods scrambled — search returns the next block over in the wrong city.
- Two different models are two different maps. "42nd Street" on the New York map and "42nd Street" on the London map are nowhere near each other. This is why querying a v1-index with v2 vectors silently fails, and why you store the model ID next to every vector.
- Dimensionality is map zoom/resolution. More dimensions can encode finer distinctions, and Matryoshka models (below) let you choose your zoom: read only the first few coordinates for a coarse "which district?", then the full list for "which door?".
- The prefixes ("query: " vs "passage: ") are the map convention: latitude-first vs longitude-first. Use the wrong convention and your coordinates land in the ocean.
Everything else in this topic — contrastive training, MTEB, cost math — is about drawing the map well and picking the right map for your city.
06.Choosing a Model: Asymmetry, Instructions, and Matryoshka
Retrieval has a structural asymmetry: queries are short questions, documents are long answers. They are different species of text that must land in the same neighborhood when relevant. Strong retrieval baselines handle this with distinct prefixes — "query: " / "passage: " (or e5's "query: "/"passage:") — so the model knows which species it is encoding. Always use the exact template the model expects; wrong prefix = wrong map convention. Instruction-tuned models (INSTRUCTOR, BGE) go further: a prompt prefix lets you steer behavior per task ("given a question, retrieve legal clauses").
Matryoshka Representation Learning (MRL) — nested dolls, the Russian stacking kind — ships in OpenAI's text-embedding-3 models (2024) and many open-source models. The trick: training arranges that the first N dimensions of a 3072-d vector are themselves a valid N-d embedding. So you can truncate to 256-d for cheap coarse recall — a 4–12x smaller, faster ANN index — then rescore the survivors with full dimensions or a second pass. It is a direct dial on your vector-DB bill (topic 163).
Long documents: most embedders are trained on windows of ≤512 tokens. Feed a model a 6000-token file and it gets averaged or truncated into a blurry "bag of averages" vector. This is a core reason RAG chunks documents (topic 166) instead of embedding whole files: keep the retrieval unit near the model's trained size.
from openai import OpenAI
import numpy as np
client = OpenAI()
def embed(texts: list[str]) -> np.ndarray:
r = client.embeddings.create(
model="text-embedding-3-small", # 1536-d, MRL-capable
input=texts,
)
v = np.array([d.embedding for d in r.data], dtype="float32")
return v / np.linalg.norm(v, axis=1, keepdims=True)
doc_vecs = embed([c.text for c in corpus_chunks]) # write-time normalize
q_vec = embed(["how do I reset my MFA device?"]) # same model, same prefix
scores = doc_vecs @ q_vec.T # cosine = inner product
top = scores[:, 0].argsort()[::-1][:5]07.Evaluating Embeddings: MTEB and, More Importantly, Your Own
MTEB (Massive Text Embedding Benchmark) is the de-facto leaderboard: roughly 50+ tasks across retrieval, reranking, clustering, classification, and bitext (translation-pair) mining. Two caveats for practitioners:
- Leaderboard rank on English retrieval ≠ your domain. Legal, medical, code, and multilingual corpora reshuffle rankings.
- If you need retrieval, compare models on the retrieval subsets only — a model can ace clustering and flop at ranking passages.
The only benchmark that decides your ship/no-ship is yours: build 100–1000 query→relevant-chunk pairs from real traffic and measure recall@k (is the right chunk inside the top-k?) with each candidate model. Expect surprises — a 768-d open model frequently beats a 3072-d API model on narrow domains.
Beyond text: multimodal embeddings (CLIP-style image↔text, and 2025 audio-capable embedders) share one space across modalities, so a text query can retrieve photos. Code retrieval needs models trained on code (jina-embeddings for code, voyage-code, Nomic), because natural-language similarity matches poorly on identifier semantics — variable names behave like rare vocabulary, not like English.
Architectural Trade-offs & Production Realities
Architectural Advantages
- One geometry serves search, clustering, dedup, recommendations, and reranking rescoring.
- MRL truncation gives a continuous quality/cost/latency dial on the same vectors.
- API embedders are cheap, fast, and good enough to launch in an afternoon.
Trade-offs & Constraints
- Silent failure on model/version mismatch; vectors are not interchangeable across maps.
- Long or mixed-language docs degrade into blurry "average" vectors.
- Raw cosine on unnormalized vectors mixes magnitude noise into ranking.
The 2024 text-embedding-3 family exposes Matryoshka-style dimension reduction (e.g., truncate 3072-d to 512-d per index) and lets customers match index size to recall targets; Azure AI Search wires these vectors into hybrid skillsets combining lexical and semantic rankers.
Staff+ Engineering Takeaways
- Retrieval embeddings are trained contrastively (InfoNCE: pull positives together, push hard negatives apart) so cosine similarity approximates human relevance.
- Use the model's expected query/passage prefix and normalize vectors; then compare with the dot product — ranking by angle is the whole game.
- Embeddings are model-versioned artifacts: two models are two incompatible maps, and swapping models requires a full re-index.
- Matryoshka (MRL) models let you truncate dimensions for cheaper coarse retrieval, then rescore at full size.
- MTEB is a starting shortlist only — validate with recall@k on your own domain eval set.
Topic Knowledge Check
Exercise 1 of 3 • Test your architectural comprehension.
Why does cosine similarity equal the dot product for most retrieval stacks?
How clear and actionable was this distributed systems breakdown?