TOPIC #217Advanced 14 min read

State Space Models: Mamba and the Post-Attention Quest

AI
AI & ML Editorial
Report an issue
Key takeawayCore Concept Summary

State space models try to get the best of both worlds: recurrent inference with a fixed-size memory (no growing KV cache) but Transformer-style parallel training. This topic follows the line from S4 to Mamba's input-dependent "selection" and hardware-aware scan, to Mamba-2's proof that "Transformers are SSMs," and then to the honest 2024 verdict — pure SSMs fail exact retrieval, so the survivors are hybrids that keep a few attention layers.

01.The Problem: A Memory Bill That Grows Forever

Recall the KV-cache story (previous topic, in one plain sentence: an LLM stores every past token's Key and Value vectors and must read that whole stack for every generated token).

That stack never stops growing.

  • 100K tokens of context ≈ 13 GB of cache on an 8B-class model.
  • Decode time per token rises with context, because the bytes read rise with context.
  • Long conversations eventually run out of GPU memory, not out of ideas.

Attention gets this exactness for free — it can quote token 47 of a 1M-token prompt word-perfect. But do we always need the whole transcript on the desk?

Human brains certainly do not: you listen to a one-hour lecture with a fixed-size head, not with a growing stack of verbatim pages.

So the question becomes

Insight

Can we get Transformer-quality language models that carry a constant-size memory — like an RNN — but still train in parallel, like a Transformer?

Classic RNNs had the right memory shape (a hidden state h updated per step, O(1) per token forever) but two fatal bugs: they train serially through time (no parallelism across the sequence) and their gradients vanish or explode over long horizons.

State space models (SSMs) are the 2021–2024 attempt to fix both at once. This topic is that story — and its humbling sequel.

Two Faces of an SSM: Recurrent Inference, Parallel Training 🧮

PRO Architecture Blueprint

Two Faces of an SSM: Recurrent Inference, Parallel Training 🧮

SSMs maintain a fixed-size hidden state updated per step (cheap constant-memory decode) while the same computation can be expressed as a global convolution/scan computed in parallel across training sequences.

Two Faces of an SSM: Recurrent Inference, Parallel Training 🧮
100%
Touchpad: Pinch to zoom • Drag to pan
Rendering visual architecture flowchart...
PRO & LIFETIME CURRICULUM

Unlock Topic #217: State Space Models: Mamba and the Post-Attention Quest

You are viewing a preview. The full in-depth technical walkthrough, worked derivations, and code notebooks for this concept, along with self-assessment quizzes, are available with Pro or Lifetime Access.

Production Deep Dive

Failure modes, high-throughput bottlenecks, and real FAANG implementation decisions.

Interactive Blueprints

Interactive system topology diagrams, live parameter simulators, and downloadable SVG charts.

Knowledge Assessment

Staff-level multiple-choice quiz questions with instant feedback and answer explanations.

Cross-Device Progress Sync

Firebase Google authentication automatically syncs your completed topics and quiz scores.

Rate This Architecture ChapterFeedback & Rating

How clear and actionable was this distributed systems breakdown?