Causal and Autoregressive Language Modeling
Every LLM rests on one idea: write the probability of a whole sequence as a chain of next-token steps, P(x) = Π P(xₜ | x<ₜ). Train all the steps at once with a triangular attention mask and teacher forcing; run them one at a time at inference with a KV cache; then decide, by your sampling rule, how bold or careful the model sounds.
01.The Problem: A Sentence Is Not a Bag of Words
Suppose a model knows single-word probabilities:
P("the") = 0.05, P("cat") = 0.01, P("banana") = 0.002
What is the probability of the whole sentence "the cat sat"?
You cannot just multiply those numbers — that would treat the words as independent, and "banana elephant Tuesday" would score like real English.
The probability of a sentence depends on order and context: P("cat") is tiny alone, but large right after "the".
So the real question is how to define and compute
P(entire sequence of tokens)
in a way a neural network can both train and sample from.
One option (which BERT effectively chose, Concept 135): ask each position "guess your own missing word from the neighbors." Great for understanding — but the result is not a probability of whole sequences, so you cannot generate with it.
The other option is the one every GPT-class model uses: insist that the sequence unfolds left to right, one genuinely conditional step at a time. That family is the causal (autoregressive) language model, and its math fits on one line.
The Autoregressive Generate Loop 🔁
The Autoregressive Generate Loop 🔁
Training computes all positions in parallel once; inference loops one token at a time, leaning on the KV cache so each step only pays for the newest position.
Unlock Topic #138: Causal and Autoregressive Language Modeling
You are viewing a preview. The full in-depth technical walkthrough, worked derivations, and code notebooks for this concept, along with self-assessment quizzes, are available with Pro or Lifetime Access.
Failure modes, high-throughput bottlenecks, and real FAANG implementation decisions.
Interactive system topology diagrams, live parameter simulators, and downloadable SVG charts.
Staff-level multiple-choice quiz questions with instant feedback and answer explanations.
Firebase Google authentication automatically syncs your completed topics and quiz scores.
How clear and actionable was this distributed systems breakdown?