The Dot Product: Similarity, Projections, and Attention
The dot product takes two lists of numbers, multiplies matching entries, and adds them up. One number out. That single recipe gives you cosine similarity, vector-database search, every neuron, and the Q·K scores that let Transformers "attend".
Dot Products Are Attention
Scaled dot-product attention: score every key against the query with a dot product, softmax the scores, then form a weighted sum (linear combination) of value vectors.
01.The Problem: Are Two Lists of Numbers "Alike"?
You now know a vector is just a list of numbers. Models love turning things into lists: a word becomes a list, a user becomes a list, a document becomes a list.
Once everything is a list, you need to answer questions like
Is this document about the same topic as that query?
Which stored picture is closest to the one I just got?
Should this neuron fire loudly or stay quiet?
All three reduce to one question
How aligned are two lists of numbers?
"Aligned" is fuzzy. You want a single number: big when the lists agree, zero when they're unrelated, negative when they push opposite ways. That number comes from one tiny operation: the dot product.
02.The Idea in Plain Words: Multiply Matching Entries, Then Add
The whole operation in one line
The dot product (inner product) is: multiply corresponding components, then sum the results.
For two vectors a, b ∈ Rⁿ:
a · b = Σᵢ aᵢbᵢ
Read the symbol · as "dot". The output is a scalar — one plain number, not a list.
Four properties that matter in ML:
- Symmetric:
a · b = b · a— alignment doesn't care which comes first. - Self product gives length:
a · a = ‖a‖²— dot a vector with itself and you get its squared length (connects to Topic 8, Norms). - Zero means orthogonal: the vectors point in independent directions; e.g., after centering, uncorrelated features have near-zero dot products.
- Matrix multiplication is just many dot products: the entry
ABbuilds from dot products between rows ofAand columns ofB(Topic 4).
import numpy as np
a = np.array([1.0, 2.0, 3.0])
b = np.array([0.5, -1.0, 2.0])
print(a @ b) # 5.5 (also np.dot / np.inner)
cos = (a @ b) / (np.linalg.norm(a) * np.linalg.norm(b))
print(cos) # ~0.798 : angle between embeddings
# Similarity of one query against 3 stored documents (vector-DB style)
docs = np.array([[1,0,1],[0,1,1],[1,1,0]], dtype=float)
query = np.array([1.0, 0.0, 0.5])
print(docs @ query) # dot product for every doc at once03.A Tiny Worked Example, By Hand
Take a = [1, 2, 3] and b = [0.5, -1, 2]. Line them up and go entry by entry.
- position 1:
1 × 0.5 = 0.5 - position 2:
2 × (-1) = -2 - position 3:
3 × 2 = 6 - now add them:
0.5 − 2 + 6 = 4.5
So a · b = 4.5. That's the entire algorithm: multiply, then add.
A second one that shows the "alignment" idea even better. Say you grade by three topics and two students have score vectors s₁ = [5, 3, 4], s₂ = [4, 4, 5]:
s₁ · s₂ = 5·4 + 3·4 + 4·5 = 20 + 12 + 20 = 52- A big dot product: the two profiles point the same way (similar strengths).
- If instead
s₂ = [4, -4, 0]:20 − 12 + 0 = 8→ small, the profiles disagree. - If
s₂ = [3, -5, 2]:15 − 15 + 8 = 8. Change one sign more and you gets₁ · s₂ = 0→ perfectly unrelated directions.
04.Visual Intuition: The Shadow Between Two Arrows
Draw each vector as an arrow from the origin. The dot product is a score of "how much do these two arrows point the same way?" — measured by the shadow one arrow casts on the other.
codesame direction partial agree perpendicular oppose ╱ ╱ ╱ │ ╲ ╱ ╱ big + ╱ medium + │ zero ╲ negative ● ● ● ● a≈b (long shadow) a,·b small angle a ⟂ b a,·b > 90°
- When the arrows line up, each one's "shadow" on the other is long → large positive dot product.
- When they're at right angles, the shadow shrinks to nothing → zero.
- When they point opposite, the shadow is negative → negative dot product.
That shadow picture is exactly what "projection" means: the projection of a onto a unit direction u is (a · u)u — the dot product a · u is the shadow length.
05.The Analogy: A Survey Scoreboard
Carry one image through the rest of this topic: a two-person survey scoreboard.
Imagine a survey with three questions, and each answer is a number. Your answers form a vector a, a friend's answers form b. The dot product a · b multiplies your answers question by question and adds them up — a single "how much did we agree?" score.
Big positive = we answered the same way → aligned. Near zero = no pattern of agreement → unrelated directions. Negative = we answered opposite ways → opposing.
Now the payoff: in ML, that scoreboard generalizes.
- The "questions" become features or embedding dimensions.
- The "answers" become weights, word-embedding values, query and key vectors.
- The "agreement score" becomes similarity, neuron activation, or attention score — all of them literally the same multiply-and-add.
06.Geometry: How Aligned Are Two Vectors?
The algebra above and the shadow picture connect through the law of cosines, which gives the geometric identity:
a · b = ‖a‖ ‖b‖ cos θ
So the dot product measures alignment: positive when the angle is under 90°, zero when perpendicular, negative when opposing. Read it as: the agreement score (a·b) equals the length of a, times the length of b, times how aligned they are (cos θ).
Normalizing both vectors — dividing each by its own length — strips the magnitudes and leaves just the alignment, turning the dot product into cosine similarity ∈ [−1, 1] — the default relevance metric for embeddings:
- Text/image retrieval with sentence-transformers, CLIP, or OpenAI embeddings ranks by cosine similarity.
- Normalized (cosine) vector indexes in FAISS, pgvector, and Milvus store unit vectors so the dot product is the similarity — one multiply-add pass per candidate.
- Recommendation systems compute user–item affinity as inner products of learned latent vectors (classic matrix-factorization view, still visible in two-tower models).
07.Why AI Cares: Neurons and Attention Are Just Dot Products
The scoreboard is secretly the engine of neural nets.
Inside a neuron. A fully-connected neuron computes z = w · x + b then an activation. The dot product is the neuron's entire "opinion": the weight vector w encodes a preferred direction in input space; the neuron fires strongly for inputs aligned with w. A layer is a batch of such dot products — matrix multiplication again.
Inside attention. In Transformers (Vaswani et al., 2017, and every 2024–2026 LLM), attention scores are dot products between learned query and key projections:
Attention(Q,K,V) = softmax(QKᵀ / √d_k) V
Follow the scoreboard metaphor: the query q is "what I'm looking for", each key kᵢ is a stored answer sheet, q · kᵢ scores how well they agree, softmax turns the scores into weights, and the output is a weighted sum of value vectors. The Mermaid diagram at the top draws exactly this loop.
The /√d_k scaling keeps dot products from saturating the softmax as dimension grows. Variants like Grouped-Query Attention (GQA) and Multi-head Latent Attention (MLA, DeepSeek-V2/V3, 2024–2025) change how many K/V heads are stored, but the scoring operation remains the humble dot product — executed by FlashAttention kernels as fused matrix multiplies.
08.In Practice: Computational Reality
One dot product costs n multiplies and n−1 adds — O(n). Trivial… until you need it billions of times:
- Retrieving from 10 million 1,536-dim embeddings ≈ 1.5×10¹⁰ multiply-adds per query → why Approximate Nearest Neighbor (ANN) indexes (HNSW, IVF-PQ) exist to skip most dot products.
- Long-context inference: attention cost grows with the number of query–key dot products (quadratic in sequence length), motivating FlashAttention tiling, sparse/sliding-window attention, and KV caching.
- Dot products are the "embarrassingly parallel" kernel GPUs and TPUs are built for: SIMD FMA (fused multiply-add) units chew through them at teraflop rates.
Architectural Trade-offs & Production Realities
Architectural Advantages
- Cheapest meaningful similarity metric: O(n) single pass, vectorizable, and GPU-friendly.
- Unifies neurons, attention scores, and retrieval ranking under one operation.
- On normalized vectors it equals cosine similarity, scale-invariant and bounded.
Trade-offs & Constraints
- Raw dot products on unnormalized embeddings conflate magnitude with relevance.
- Cosine similarity is blind to non-semantic norm differences that sometimes matter (recency, popularity).
- Exact dot-product search over large corpora is linear in corpus size; indexes trade accuracy for speed.
Every generated token computes dot products between its query vector and cached key vectors for up to 128K context positions (QKᵀ), softmax-scales them, and takes a weighted sum of value vectors. FlashAttention fuses these dot products into tiled GPU kernels; MLA compresses the keys that must be dot-producted, slashing memory bandwidth.
Staff+ Engineering Takeaways
- Dot product = sum of element-wise products; the result is a scalar measuring alignment.
- Geometrically a · b = ‖a‖‖b‖cos θ; normalize inputs and it becomes cosine similarity in [-1, 1].
- A neuron is w · x + b; a layer is many dot products; attention scores are learned dot products QKᵀ/√d.
- Vector databases rank candidates with dot products; ANN indexes (HNSW, IVF) exist because exact scoring is linear in corpus size.
- a · a = ‖a‖² links the dot product to norms, and row-column dot products define matrix multiplication.
Topic Knowledge Check
Exercise 1 of 3 • Test your architectural comprehension.
Vectors a and b have dot product exactly 0. What does this imply?
How clear and actionable was this distributed systems breakdown?