Modern Architecture Advances
What actually runs in 2024-2026 frontier systems:
All Topics in Phase 13
0 of 11 completedHow sparse Mixture-of-Experts architectures decouple total parameters from active parameters: a learned router sends each token to only the top-k expert FFNs, so capacity grows while per-token compute stays small — the design behind Mixtral, DeepSeek-V3, and Kimi K2.
Beating the O(N²) attention tax at long context: if each query only reads a small subset of keys, cost drops toward linear. Fixed patterns (sliding window, dilated, block-sparse), the 2025 wave of trainable sparse attention (DeepSeek Sparse Attention, Native Sparse Attention), and inference-only KV-selection sparsity.
The IO-aware exact attention kernel that made long context affordable: attention is reorganized so the N² score matrix never travels to slow GPU memory — tiles are processed on-chip with an online softmax. Plus the FA1 → FA2 → FA3 evolution on Ampere, Hopper, and Blackwell GPUs.
How rotating query and key vectors by position-encoded angles gives attention a built-in relative-position bias with zero added parameters — and why tuning the rotation base θ became the main lever for 128K-1M context extension.
How shrinking the number of K/V heads — down to one (MQA) or a few groups (GQA) — slashes KV-cache size and decode memory bandwidth at near-zero quality cost, and why GQA became the default in Llama, Mistral, Gemma, and Qwen.
An LLM writes one token at a time, and every new token must "look back" at all the previous ones. The KV cache stores each past token's Key and Value vectors once so they are never recomputed — and it becomes the single biggest consumer of GPU memory and bandwidth when serving. This topic is the story of taming that stack: PagedAttention, prefix caching, quantization, eviction, and cross-node offloading.
LLMs generate one token per step, and each step leaves most of the GPU idle. Speculative decoding fixes the wait without changing the words: a cheap "draft" model guesses several tokens ahead, the real model verifies them all in one parallel pass, and a rejection-sampling rule guarantees the output is statistically identical to the big model alone. Latency drops 2-3×; quality does not move at all.
Between 2024 and 2026, frontier models went from 4K-token windows to advertised 1M-10M windows. Growing a window is never one trick: positional encodings, the attention budget, and long-document training data must move together — and benchmarks like RULER and NoLiMa show the useful context still trails the advertised one. This topic covers how the race worked, what it costs, and how to use big windows honestly.
State space models try to get the best of both worlds: recurrent inference with a fixed-size memory (no growing KV cache) but Transformer-style parallel training. This topic follows the line from S4 to Mamba's input-dependent "selection" and hardware-aware scan, to Mamba-2's proof that "Transformers are SSMs," and then to the honest 2024 verdict — pure SSMs fail exact retrieval, so the survivors are hybrids that keep a few attention layers.
Reasoning models are trained — with reinforcement learning over long chains of thought — to write a private draft before answering: plan, check, backtrack, try again. OpenAI's o1 (Sept 2024) proved the class; DeepSeek-R1 (Jan 2025) reproduced it openly using GRPO and rewards that an auto-grader can verify, and showed the deliberation emerges even without any worked examples. This topic covers the recipe, the emergent "aha" behaviors, the distillation wave, and the very real failure modes.
For twenty years, AI capability grew by spending compute during training. From 2024-2026, a second axis matters just as much: compute spent AT INFERENCE — thinking longer (sequential) or trying more candidates and verifying them (parallel). This topic covers the scaling laws behind both, why verifier quality is the ceiling, how difficulty-based allocation lets a small model beat a 14x larger one, and why every API now sells "thinking effort" as a priced knob.