PHASE 6 CURRICULUM

Sequence Models & Attention

Progress0 of 16 (0%)

Language and time-series data arrive as sequences.

Key Architectural Domains & Syllabus
This phase walks through one-hot encodings
word embeddings
RNNs and their gradient pathologies
GRUs and LSTMs
bidirectional models
encoder-decoder seq2seq with attention
then builds the scaled dot-product and multi-head self-attention machinery that powers every modern large model
16 In-Depth Topics ~128 Minutes Reading Time Interactive Quizzes & Assessments

All Topics in Phase 6

0 of 16 completed

Sequence data is data where the order of elements is part of the meaning: "dog bites man" is not "man bites dog". This topic shows how models eat sequences — tokens, embeddings, padding masks, task shapes — and poses the hard problem that drives all of Phase 6: how do distant steps influence each other?

11 min read•3 Quiz Questions

An RNN is one small cell wired to itself: its output at step t becomes an input at step t+1, so it carries memory forward. The same weights are reused at every step — that single trick buys variable-length inputs and compact models, and costs a very long gradient path.

11 min read•3 Quiz Questions

The hidden state h_t is one fixed-size vector — the model’s entire memory of everything it has seen. Learn what a few hundred floats can store, what they provably cannot, and why the "summarize vs re-read" choice still decides 2024-2026 serving economics.

11 min read•3 Quiz Questions

To learn from a long sequence, the gradient must multiply one factor per time step. If each factor shrinks the signal, early tokens get no credit (vanishing); if each grows it, the update detonates (exploding). Follow the product with small numbers, then see which fixes actually worked.

12 min read•3 Quiz Questions

The LSTM adds a protected ledger (the cell state C_t) plus three learned gatekeepers that erase, write, and read it. Because the ledger updates by addition with gate values near 1, the gradient product stops decaying — the vanishing-gradient problem of Topic 113 gets its first real cure.

12 min read•3 Quiz Questions
#115GRU: A Leaner Gate DesignIntermediateFREE

The Gated Recurrent Unit asks: does the LSTM really need two state vectors and three gate matrices? It fuses the ledger and the readout into one vector with just two gates — 25% fewer recurrent parameters, the same gradient cure, and usually statistically tied results.

11 min read•3 Quiz Questions

When the input and output are sequences of *different lengths*, one network reads and compresses, the other expands and generates. Understand the RNN encoder-decoder pattern, the one-vector "memo" it forced between them, why that memo was a bottleneck — and why T5, BART, and Whisper still run this exact contract today.

11 min read•3 Quiz Questions

Seq2seq is the encoder-decoder pattern (Topic 116) industrialized: the 2014 deep-LSTM translation machine, teacher forcing for stable training, exposure bias as its shadow, and beam search as the inference contract. These mechanics still run every generation system — including 2026 LLM servers.

12 min read•3 Quiz Questions

The old machine-translation models squeezed a whole sentence into one fixed vector before writing the translation — and long sentences got mangled. Bahdanau attention says: keep every encoder state, and at each decoding step take a learned, softmax-weighted look-back at all of them. Score → softmax → weighted sum: that pipeline is the ancestor of all cross-attention.

12 min read•4 Quiz Questions

Take Bahdanau's machinery and point the queries at the very same sequence being read — and suddenly one layer connects any two positions directly. Here is the attention matrix, the causal mask, the O(T²) bill, and why this single primitive now runs the world's models.

13 min read•4 Quiz Questions

softmax(QKᵀ / √d_k) V is the entire mechanism of modern AI. Read it as a five-step dataflow, see with tiny numbers exactly why the √d_k divisor is load-bearing (unscaled scores saturate softmax and kill gradients), and trace the formula to the kernels and KV caches that run it in 2026.

13 min read•4 Quiz Questions

Attention is best understood as a soft lookup into a memory: a query asks, keys advertise, values deliver. Here is what W_Q, W_K, W_V actually learn, why queries and keys must be separate tensors, and how this one viewpoint explains KV caches, retrieval-augmented generation, and 2024-2026 memory research.

13 min read•4 Quiz Questions

One attention matrix per layer is a rank bottleneck: a single softmax cannot attend to syntax and coreference at the same time without blurring them. Multi-head attention runs h smaller attentions in different learned subspaces and merges them — and its parameterization became the 2024-2026 KV-cache battleground: MHA → MQA → GQA → MLA.

13 min read•4 Quiz Questions

Self-attention is blind to order — it reads a sentence as a bag of vectors — so position must be injected into the representation. From sinusoids and learned embeddings to the 2024-2026 regime of RoPE, ALiBi, and context-extension tricks that decide how far a model can actually see.

13 min read•4 Quiz Questions

You have the parts — attention, heads, positions. Now assemble the machine: stacked attention + FFN with residuals and LayerNorm, the encoder/decoder/cross-attention wiring, and the three model families that split in 2018 — with decoder-only winning 2020-2026.

14 min read•4 Quiz Questions

After attention mixes tokens, a two-layer MLP processes each position alone — and it holds roughly two-thirds of every Transformer's parameters. What the FFN actually computes, why it grew from 4× ReLU to SwiGLU and MoE, and what 2024-2026 research thinks it stores.

14 min read•4 Quiz Questions