Multi-Head Attention: Parallel Subspaces of Look
One attention matrix per layer is a rank bottleneck: a single softmax cannot attend to syntax and coreference at the same time without blurring them. Multi-head attention runs h smaller attentions in different learned subspaces and merges them — and its parameterization became the 2024-2026 KV-cache battleground: MHA → MQA → GQA → MLA.
01.The Problem: One Softmax Cannot Ask Two Questions at Once
You now know the recipe: score, softmax, blend (Topics 119-121).
But look closely at softmax — it hides a tyranny.
Suppose the verb "gave" needs two different lookups in the same sentence:
- the subject (for grammatical agreement),
- the object (for meaning: what was given).
A single attention row must put its weights on one shared distribution.
- Attend fully to the subject? Then the object gets ~0.
- Split the difference?
α = (0.5, 0.5)— a blur of both retrievals mixed into one vector.
You cannot run two separate lookups on one ballot.
Softmax forces competition: a query that must attend to both "the subject (for agreement)" and "the object (for semantics)" can only split mass, blurring two distinct retrievals into one mixture.
And the problem compounds across the layer: one q · k bilinear form per layer is a low-rank similarity — the model wants many different relevance notions (syntax, reference, proximity, semantics) simultaneously.
The fix is embarrassingly simple:
Do not vote once. Sit on a panel.
Run h small attention layers in parallel, each with its own learned notion of "relevant," then merge their answers.
That is multi-head attention (Vaswani §3.2.2): it gives each notion its own subspace.
Split, Attend in Parallel, Re-merge 🔀
Split, Attend in Parallel, Re-merge 🔀
Each head is a full attention over a 64-dim learned subspace. Total cost is essentially unchanged from one 512-dim head — 8 narrow looks instead of one wide one — and the narrowness is the point.
Unlock Topic #122: Multi-Head Attention: Parallel Subspaces of Look
You are viewing a preview. The full in-depth technical walkthrough, worked derivations, and code notebooks for this concept, along with self-assessment quizzes, are available with Pro or Lifetime Access.
Failure modes, high-throughput bottlenecks, and real FAANG implementation decisions.
Interactive system topology diagrams, live parameter simulators, and downloadable SVG charts.
Staff-level multiple-choice quiz questions with instant feedback and answer explanations.
Firebase Google authentication automatically syncs your completed topics and quiz scores.
How clear and actionable was this distributed systems breakdown?