PHASE 13 CURRICULUM

Modern Architecture Advances

Progress0 of 11 (0%)

What actually runs in 2024-2026 frontier systems:

Key Architectural Domains & Syllabus
sparse mixture-of-experts routing
FlashAttention and sparse attention kernels
rotary position embeddings and length extrapolation
grouped-query and multi-query attention for cheap inference
KV-cache management
speculative decoding
long-context training and evaluation
state-space models like Mamba
the test-time-compute scaling of reasoning models
11 In-Depth Topics ~88 Minutes Reading Time Interactive Quizzes & Assessments

All Topics in Phase 13

0 of 11 completed
#209Mixture of Experts (MoE)AdvancedFREE

How sparse Mixture-of-Experts architectures decouple total parameters from active parameters: a learned router sends each token to only the top-k expert FFNs, so capacity grows while per-token compute stays small — the design behind Mixtral, DeepSeek-V3, and Kimi K2.

14 min read•3 Quiz Questions
#210Sparse AttentionAdvancedFREE

Beating the O(N²) attention tax at long context: if each query only reads a small subset of keys, cost drops toward linear. Fixed patterns (sliding window, dilated, block-sparse), the 2025 wave of trainable sparse attention (DeepSeek Sparse Attention, Native Sparse Attention), and inference-only KV-selection sparsity.

14 min read•3 Quiz Questions
#211Flash AttentionAdvancedFREE

The IO-aware exact attention kernel that made long context affordable: attention is reorganized so the N² score matrix never travels to slow GPU memory — tiles are processed on-chip with an online softmax. Plus the FA1 → FA2 → FA3 evolution on Ampere, Hopper, and Blackwell GPUs.

14 min read•3 Quiz Questions

How rotating query and key vectors by position-encoded angles gives attention a built-in relative-position bias with zero added parameters — and why tuning the rotation base θ became the main lever for 128K-1M context extension.

14 min read•3 Quiz Questions

How shrinking the number of K/V heads — down to one (MQA) or a few groups (GQA) — slashes KV-cache size and decode memory bandwidth at near-zero quality cost, and why GQA became the default in Llama, Mistral, Gemma, and Qwen.

14 min read•3 Quiz Questions

An LLM writes one token at a time, and every new token must "look back" at all the previous ones. The KV cache stores each past token's Key and Value vectors once so they are never recomputed — and it becomes the single biggest consumer of GPU memory and bandwidth when serving. This topic is the story of taming that stack: PagedAttention, prefix caching, quantization, eviction, and cross-node offloading.

14 min read•3 Quiz Questions
#215Speculative DecodingAdvanced PRO

LLMs generate one token per step, and each step leaves most of the GPU idle. Speculative decoding fixes the wait without changing the words: a cheap "draft" model guesses several tokens ahead, the real model verifies them all in one parallel pass, and a rejection-sampling rule guarantees the output is statistically identical to the big model alone. Latency drops 2-3×; quality does not move at all.

14 min read•4 Quiz Questions

Between 2024 and 2026, frontier models went from 4K-token windows to advertised 1M-10M windows. Growing a window is never one trick: positional encodings, the attention budget, and long-document training data must move together — and benchmarks like RULER and NoLiMa show the useful context still trails the advertised one. This topic covers how the race worked, what it costs, and how to use big windows honestly.

14 min read•3 Quiz Questions

State space models try to get the best of both worlds: recurrent inference with a fixed-size memory (no growing KV cache) but Transformer-style parallel training. This topic follows the line from S4 to Mamba's input-dependent "selection" and hardware-aware scan, to Mamba-2's proof that "Transformers are SSMs," and then to the honest 2024 verdict — pure SSMs fail exact retrieval, so the survivors are hybrids that keep a few attention layers.

14 min read•4 Quiz Questions

Reasoning models are trained — with reinforcement learning over long chains of thought — to write a private draft before answering: plan, check, backtrack, try again. OpenAI's o1 (Sept 2024) proved the class; DeepSeek-R1 (Jan 2025) reproduced it openly using GRPO and rewards that an auto-grader can verify, and showed the deliberation emerges even without any worked examples. This topic covers the recipe, the emergent "aha" behaviors, the distillation wave, and the very real failure modes.

14 min read•4 Quiz Questions

For twenty years, AI capability grew by spending compute during training. From 2024-2026, a second axis matters just as much: compute spent AT INFERENCE — thinking longer (sequential) or trying more candidates and verifying them (parallel). This topic covers the scaling laws behind both, why verifier quality is the ceiling, how difficulty-based allocation lets a small model beat a 14x larger one, and why every API now sells "thinking effort" as a priced knob.

14 min read•4 Quiz Questions