State Space Models: Mamba and the Post-Attention Quest
State space models try to get the best of both worlds: recurrent inference with a fixed-size memory (no growing KV cache) but Transformer-style parallel training. This topic follows the line from S4 to Mamba's input-dependent "selection" and hardware-aware scan, to Mamba-2's proof that "Transformers are SSMs," and then to the honest 2024 verdict — pure SSMs fail exact retrieval, so the survivors are hybrids that keep a few attention layers.
01.The Problem: A Memory Bill That Grows Forever
Recall the KV-cache story (previous topic, in one plain sentence: an LLM stores every past token's Key and Value vectors and must read that whole stack for every generated token).
That stack never stops growing.
- 100K tokens of context ≈ 13 GB of cache on an 8B-class model.
- Decode time per token rises with context, because the bytes read rise with context.
- Long conversations eventually run out of GPU memory, not out of ideas.
Attention gets this exactness for free — it can quote token 47 of a 1M-token prompt word-perfect. But do we always need the whole transcript on the desk?
Human brains certainly do not: you listen to a one-hour lecture with a fixed-size head, not with a growing stack of verbatim pages.
So the question becomes
Can we get Transformer-quality language models that carry a constant-size memory — like an RNN — but still train in parallel, like a Transformer?
Classic RNNs had the right memory shape (a hidden state h updated per step, O(1) per token forever) but two fatal bugs: they train serially through time (no parallelism across the sequence) and their gradients vanish or explode over long horizons.
State space models (SSMs) are the 2021–2024 attempt to fix both at once. This topic is that story — and its humbling sequel.
Two Faces of an SSM: Recurrent Inference, Parallel Training 🧮
Two Faces of an SSM: Recurrent Inference, Parallel Training 🧮
SSMs maintain a fixed-size hidden state updated per step (cheap constant-memory decode) while the same computation can be expressed as a global convolution/scan computed in parallel across training sequences.
Unlock Topic #217: State Space Models: Mamba and the Post-Attention Quest
You are viewing a preview. The full in-depth technical walkthrough, worked derivations, and code notebooks for this concept, along with self-assessment quizzes, are available with Pro or Lifetime Access.
Failure modes, high-throughput bottlenecks, and real FAANG implementation decisions.
Interactive system topology diagrams, live parameter simulators, and downloadable SVG charts.
Staff-level multiple-choice quiz questions with instant feedback and answer explanations.
Firebase Google authentication automatically syncs your completed topics and quiz scores.
How clear and actionable was this distributed systems breakdown?