TOPIC #122Advanced 13 min read

Multi-Head Attention: Parallel Subspaces of Look

AI
AI & ML Editorial
Report an issue
Key takeawayCore Concept Summary

One attention matrix per layer is a rank bottleneck: a single softmax cannot attend to syntax and coreference at the same time without blurring them. Multi-head attention runs h smaller attentions in different learned subspaces and merges them — and its parameterization became the 2024-2026 KV-cache battleground: MHA → MQA → GQA → MLA.

01.The Problem: One Softmax Cannot Ask Two Questions at Once

You now know the recipe: score, softmax, blend (Topics 119-121).

But look closely at softmax — it hides a tyranny.

Suppose the verb "gave" needs two different lookups in the same sentence:

  • the subject (for grammatical agreement),
  • the object (for meaning: what was given).

A single attention row must put its weights on one shared distribution.

  • Attend fully to the subject? Then the object gets ~0.
  • Split the difference? α = (0.5, 0.5) — a blur of both retrievals mixed into one vector.
Insight

You cannot run two separate lookups on one ballot.

Softmax forces competition: a query that must attend to both "the subject (for agreement)" and "the object (for semantics)" can only split mass, blurring two distinct retrievals into one mixture.

And the problem compounds across the layer: one q · k bilinear form per layer is a low-rank similarity — the model wants many different relevance notions (syntax, reference, proximity, semantics) simultaneously.

The fix is embarrassingly simple:

Insight

Do not vote once. Sit on a panel.

Run h small attention layers in parallel, each with its own learned notion of "relevant," then merge their answers.

That is multi-head attention (Vaswani §3.2.2): it gives each notion its own subspace.

Split, Attend in Parallel, Re-merge 🔀

PRO Architecture Blueprint

Split, Attend in Parallel, Re-merge 🔀

Each head is a full attention over a 64-dim learned subspace. Total cost is essentially unchanged from one 512-dim head — 8 narrow looks instead of one wide one — and the narrowness is the point.

Split, Attend in Parallel, Re-merge 🔀
100%
Touchpad: Pinch to zoom • Drag to pan
Rendering visual architecture flowchart...
PRO & LIFETIME CURRICULUM

Unlock Topic #122: Multi-Head Attention: Parallel Subspaces of Look

You are viewing a preview. The full in-depth technical walkthrough, worked derivations, and code notebooks for this concept, along with self-assessment quizzes, are available with Pro or Lifetime Access.

Production Deep Dive

Failure modes, high-throughput bottlenecks, and real FAANG implementation decisions.

Interactive Blueprints

Interactive system topology diagrams, live parameter simulators, and downloadable SVG charts.

Knowledge Assessment

Staff-level multiple-choice quiz questions with instant feedback and answer explanations.

Cross-Device Progress Sync

Firebase Google authentication automatically syncs your completed topics and quiz scores.

Rate This Architecture ChapterFeedback & Rating

How clear and actionable was this distributed systems breakdown?