PHASE 8 CURRICULUM

Large Language Models

Progress0 of 20 (0%)

Understand the models everyone now builds on:

Key Architectural Domains & Syllabus
the decoder-only transformer architecture
causal pretraining objectives
datasets and tokenizers at trillion-token scale
scaling laws (Kaplan and Chinchilla)
emergence and in-context learning
instruction tuning and RLHF overviews
decoding strategies (temperature, top-k/p, beam)
prompt engineering and chain-of-thought
context windows
the embedding APIs that anchor semantic search
20 In-Depth Topics ~160 Minutes Reading Time Interactive Quizzes & Assessments

All Topics in Phase 8

0 of 20 completed

How base language models are built: one self-supervised game — guess the next token — replayed on trillions of tokens of curated web text. This topic covers the objective, the tokenizer and corpus engineering that bound what a fixed budget can learn, the training dynamics of thousand-GPU runs, and why the model that comes out still cannot follow instructions.

14 min read•3 Quiz Questions

Training a model costs a fixed pile of compute — how do you split it between parameters and data? Scaling laws say loss falls in smooth, predictable power laws, so you can measure small and forecast the $100M run. Chinchilla added the correction: spend on both equally, roughly 20 tokens per parameter — and production models then overtrain small ones on purpose because serving, not training, dominates lifetime cost.

14 min read•2 Quiz Questions

Some AI abilities look like they switch on overnight once a model is big enough — like water suddenly boiling. A famous 2023 critique says the "switch" is often just your grading rule: the skill was growing smoothly the whole time, but your test only reports all-or-nothing scores. This topic walks through both sides and what you should actually do when planning products.

12 min read•2 Quiz Questions

A base model can finish any text, but it will not answer your question — it has never been shown what an "answer" looks like. Supervised fine-tuning (SFT) fixes that with thousands of (instruction, ideal-response) examples, training only on the response part. This topic covers what actually changes, why a small high-quality dataset (LIMA) rivals a huge one, and the limits imitation leaves behind.

12 min read•3 Quiz Questions

SFT can copy ideal answers, but "helpful" is a ranking between many good answers, not one correct string. RLHF learns that ranking from human comparisons: fine-tune, collect preference votes, distill them into a reward model, then optimize with PPO on a KL leash. This is the pipeline that became ChatGPT — plus its costs, failure modes, and the cheaper alternatives it spawned.

13 min read•3 Quiz Questions

Humans cannot rank a million answers a day, so RLHF distills their votes into one cheap scorer: a reward model trained on pairwise comparisons with the Bradley–Terry loss. This topic shows how rankings become a number, why only differences between numbers mean anything, why reward models surprisingly generalize, and where over-optimization and reward hacking put a hard ceiling on them.

12 min read•2 Quiz Questions

Reinforcement learning can destroy a language model with one greedy update, so PPO leashes every step: it clips how far a single update may move each token's probability, adds a KL leash to the old model, and uses a critic to spread one final reward across thousands of tokens. This topic works the clip through with numbers, maps it onto language generation, and explains why GRPO and DPO now challenge PPO.

13 min read•3 Quiz Questions

RLHF needs a reward model, a critic, rollouts, and four models in memory. DPO proves you can skip all of it: the reward RLHF would have learned is already implicit in the policy itself, so preference pairs can be learned with one plain classification loss. This topic derives that trick with tiny numbers, then shows exactly when DPO still loses to real RL.

12 min read•2 Quiz Questions

Making an AI harmless the classic way means paid humans reading the worst content the model can produce. Constitutional AI flips that: write the rules down as a public document called a constitution, then let the model critique and rewrite its own answers against those rules, and let an AI judge pick better responses. This topic covers the two stages (SL-CAI and RL-CAI), what a constitution contains, and where the method wins and caps out.

15 min read•2 Quiz Questions

Every chat assistant has one message the user never wrote: the system prompt, which sets persona, policies, and tool rules. Its authority comes from training, not architecture — which is exactly why "ignore your previous instructions" ever worked, and why prompt injection (OWASP LLM01) makes this channel a security boundary. This topic covers what the channel is, how production prompts are structured, and how to defend them.

15 min read•3 Quiz Questions

A "200K token window" sounds like a free bookshelf, but length is paid for twice — quadratic attention at prefill and a huge KV cache at decode — and models quietly read the middle of long inputs far worse than the edges. This topic explains why windows exist, how RoPE tricks like Positional Interpolation and YaRN extend them cheaply, and what context engineering actually means in 2025–2026.

16 min read•4 Quiz Questions

Temperature is one dial placed between the model's scores and its word choice. Low temperature squeezes the outcome lottery toward the top pick — great for facts, code, and JSON. High temperature flattens the odds — creative, but error-prone. Here is what the knob actually touches, where it misleads, and the 2024–2026 findings on when temperature matters (and when it silently hurts reasoning models).

11 min read•2 Quiz Questions

Left alone, an LLM's word lottery eventually picks something silly — the probability "tail" is huge and full of nonsense. Top-k chops off everything except the best k candidates; nucleus (top-p) keeps only the smallest group of words that adds up to p of the probability. This topic covers why unbounded sampling degenerates, why the fixed cutoff of top-k fails across contexts, and why nucleus became temperature's default companion.

11 min read•2 Quiz Questions

Once you can score each next word, you have to decide how to build a whole sentence. Three philosophies compete: greedy decoding (take the best next word now), beam search (keep the best B partial sentences and prune), and sampling (roll the probability dice). This topic explains all three — why beams ruled 2016-era machine translation, why they fail at open-ended chat, and why search is now returning inside agents.

12 min read•3 Quiz Questions

An LLM was never trained to tell the truth — it was trained to sound right. That single gap explains why fluent false case law, dead URLs, and phantom APIs are the default failure mode, not a glitch. Here is the mechanism, the intrinsic/extrinsic taxonomy, how we measure it (TruthfulQA and beyond), and the mitigation stack that actually moves the numbers.

12 min read•3 Quiz Questions

A model answering from memory will confidently invent what it does not know. Grounding is the engineering answer: first fetch real evidence, then require the model to answer only from it, with citations a program can verify. This topic covers RAG, citation contracts, attribution checks, tool-based fact-checking, and the 2024–2026 shift from static RAG to agentic search.

12 min read•3 Quiz Questions

You cannot edit a hosted LLM's weights — the only thing you control is the text you send. That makes the prompt a program, written in English. This topic covers the anatomy of a production prompt, the technique taxonomy (zero-shot, few-shot, chain-of-thought, decomposition), why evals instead of vibes decide every edit, and the 2025 reframing into context engineering.

12 min read•3 Quiz Questions

There are three ways to hand a task to a frozen model: describe it (zero-shot), show one worked example (one-shot), or show a handful (few-shot). This topic covers the GPT-3 origin of the idea, the surprising research on what examples actually teach (format as much as labels), the sensitivities that quietly break production prompts (order, label balance, selection), and a 2024–2026 decision playbook.

12 min read•3 Quiz Questions

Few-shot prompting works — but the weights never moved, so what exactly is "learning"? In-context learning is the answer: a frozen Transformer matches your new input against the examples in its window and adapts entirely through activations. This topic covers the precise definition, the induction-head circuit that explains much of it, the competing research stories (implicit gradient descent, Bayesian task inference), the hard limits, and how ICL differs from fine-tuning.

12 min read•3 Quiz Questions

Ask a 2022 LLM a two-step math problem and it blurts a wrong answer in one gulp. Chain-of-thought prompting fixed that by making the model write intermediate steps — the "show your work" trick that jumped GSM8K from 17.8% to 56.9%. This topic covers the discovery, the toolkit (few-shot/zero-shot CoT, self-consistency, verification), the faithfulness caveat, and how prompted CoT grew into RL-trained reasoning models (o1, DeepSeek-R1) that internalize the scratchpad.

14 min read•4 Quiz Questions