Backpropagation: The Chain Rule, Engineered
A network has 100 million weights; training needs a gradient for every one of them. Brute force would take 100 million forward passes. Backprop gets all of them for about twice the cost of one — by running the chain rule backward through the cached forward graph. Build the δ recursion, split the two gradients, and see what autodiff actually is.
01.The Problem: 100 Million Knobs, and You Need Them All
You know the forward pass (previous topic): numbers flow input → layers → prediction. And you know the goal of training (gradient topics): nudge every weight a tiny bit downhill on the loss, using θ ← θ − lr·∇L.
So the question becomes
How do I compute ∂L/∂w for every single weight?
For a network with 10^8 parameters, that's one hundred million partial derivatives. Per training step. Before backprop existed, you had two ideas, and both are disasters:
- Numerical differentiation (brute force): nudge one weight, re-run the forward pass, see how the loss moved. That's ~2 forward passes per parameter — about
10^8forward passes per step. A single update would take longer than the heat death of the universe. Absurd. - Symbolic differentiation: ask math software to write one giant closed-form expression for
∂L/∂wacross a 50-layer net. The expression explodes combinatorially, and every shared sub-computation gets rewritten and recomputed dozens of times.
The insight that saved deep learning: the network is a composition of simple scalar/tensor ops, each with a trivial local derivative. If you exploit that graph structure, you get all gradients in about 2× the cost of one forward pass — no matter how many parameters share it.
Historically the algorithm was published for neural nets by Linnainmaa (1970), rediscovered by Werbos (1974), and popularized by Rumelhart, Hinton & Williams (1986). It is called backpropagation.
Forward Tape, Backward Replay
Forward Tape, Backward Replay
Backprop is reverse-mode autodiff: forward caches every op's inputs; backward walks the graph in reverse, multiplying local Jacobians — gradients for weights reuse the cached activations.
Unlock Topic #83: Backpropagation: The Chain Rule, Engineered
You are viewing a preview. The full in-depth technical walkthrough, worked derivations, and code notebooks for this concept, along with self-assessment quizzes, are available with Pro or Lifetime Access.
Failure modes, high-throughput bottlenecks, and real FAANG implementation decisions.
Interactive system topology diagrams, live parameter simulators, and downloadable SVG charts.
Staff-level multiple-choice quiz questions with instant feedback and answer explanations.
Firebase Google authentication automatically syncs your completed topics and quiz scores.
How clear and actionable was this distributed systems breakdown?