Sequence Models & Attention
Language and time-series data arrive as sequences.
All Topics in Phase 6
0 of 16 completedSequence data is data where the order of elements is part of the meaning: "dog bites man" is not "man bites dog". This topic shows how models eat sequences — tokens, embeddings, padding masks, task shapes — and poses the hard problem that drives all of Phase 6: how do distant steps influence each other?
An RNN is one small cell wired to itself: its output at step t becomes an input at step t+1, so it carries memory forward. The same weights are reused at every step — that single trick buys variable-length inputs and compact models, and costs a very long gradient path.
The hidden state h_t is one fixed-size vector — the model’s entire memory of everything it has seen. Learn what a few hundred floats can store, what they provably cannot, and why the "summarize vs re-read" choice still decides 2024-2026 serving economics.
To learn from a long sequence, the gradient must multiply one factor per time step. If each factor shrinks the signal, early tokens get no credit (vanishing); if each grows it, the update detonates (exploding). Follow the product with small numbers, then see which fixes actually worked.
The LSTM adds a protected ledger (the cell state C_t) plus three learned gatekeepers that erase, write, and read it. Because the ledger updates by addition with gate values near 1, the gradient product stops decaying — the vanishing-gradient problem of Topic 113 gets its first real cure.
The Gated Recurrent Unit asks: does the LSTM really need two state vectors and three gate matrices? It fuses the ledger and the readout into one vector with just two gates — 25% fewer recurrent parameters, the same gradient cure, and usually statistically tied results.
When the input and output are sequences of *different lengths*, one network reads and compresses, the other expands and generates. Understand the RNN encoder-decoder pattern, the one-vector "memo" it forced between them, why that memo was a bottleneck — and why T5, BART, and Whisper still run this exact contract today.
Seq2seq is the encoder-decoder pattern (Topic 116) industrialized: the 2014 deep-LSTM translation machine, teacher forcing for stable training, exposure bias as its shadow, and beam search as the inference contract. These mechanics still run every generation system — including 2026 LLM servers.
The old machine-translation models squeezed a whole sentence into one fixed vector before writing the translation — and long sentences got mangled. Bahdanau attention says: keep every encoder state, and at each decoding step take a learned, softmax-weighted look-back at all of them. Score → softmax → weighted sum: that pipeline is the ancestor of all cross-attention.
Take Bahdanau's machinery and point the queries at the very same sequence being read — and suddenly one layer connects any two positions directly. Here is the attention matrix, the causal mask, the O(T²) bill, and why this single primitive now runs the world's models.
softmax(QKᵀ / √d_k) V is the entire mechanism of modern AI. Read it as a five-step dataflow, see with tiny numbers exactly why the √d_k divisor is load-bearing (unscaled scores saturate softmax and kill gradients), and trace the formula to the kernels and KV caches that run it in 2026.
Attention is best understood as a soft lookup into a memory: a query asks, keys advertise, values deliver. Here is what W_Q, W_K, W_V actually learn, why queries and keys must be separate tensors, and how this one viewpoint explains KV caches, retrieval-augmented generation, and 2024-2026 memory research.
One attention matrix per layer is a rank bottleneck: a single softmax cannot attend to syntax and coreference at the same time without blurring them. Multi-head attention runs h smaller attentions in different learned subspaces and merges them — and its parameterization became the 2024-2026 KV-cache battleground: MHA → MQA → GQA → MLA.
Self-attention is blind to order — it reads a sentence as a bag of vectors — so position must be injected into the representation. From sinusoids and learned embeddings to the 2024-2026 regime of RoPE, ALiBi, and context-extension tricks that decide how far a model can actually see.
You have the parts — attention, heads, positions. Now assemble the machine: stacked attention + FFN with residuals and LayerNorm, the encoder/decoder/cross-attention wiring, and the three model families that split in 2018 — with decoder-only winning 2020-2026.