Adam: Adaptive Moment Estimation
Adam is the default optimizer of the deep-learning era because it answers a question momentum never could: how big should THIS knob's step be, given how this knob has behaved in the past? It does so with two running averages — of the gradients (momentum) and of the squared gradients (a per-knob size gauge) — plus a small correction for starting the averages at zero. This topic walks every symbol, the AdamW weight-decay fix, and the memory cost that shapes GPU clusters.
01.The Problem: One Learning Rate, Wildly Different Knobs
Recall the setup. A neural net has millions of knobs (parameters), and gradient descent turns them all with one global learning rate η.
Now look at a real transformer's knobs:
- an embedding row for a rare word — updated maybe twice a batch; when it moves, its gradient is huge.
- a LayerNorm scale — updated for every token; its gradients are tiny and steady.
- millions of things in between.
Suppose knob A gets gradients around 10.0 and knob B gets gradients around 0.0001.
One η cannot be right for both.
- η sized for A leaves B essentially frozen.
- η sized for B makes A lurch violently every time the rare word appears (hello, exploding-update territory — Topic 86).
And momentum (Topic 88) does NOT fix this. Its velocity v is built from the same gradients, so it just reproduces the 100,000× size gap in a smoothed form.
So the question becomes
What if each knob carried its own size gauge, and we set its step in units of its own recent history?
That idea — plus momentum, plus one small bookkeeping fix — is Adam. It has been the default optimizer for essentially every language model since 2015.
Adam Dataflow
Adam Dataflow
Two exponential moving averages — of g and of g² — normalized by their bias-correction factors; the step is momentum direction divided by per-coordinate RMS magnitude.
Unlock Topic #89: Adam: Adaptive Moment Estimation
You are viewing a preview. The full in-depth technical walkthrough, worked derivations, and code notebooks for this concept, along with self-assessment quizzes, are available with Pro or Lifetime Access.
Failure modes, high-throughput bottlenecks, and real FAANG implementation decisions.
Interactive system topology diagrams, live parameter simulators, and downloadable SVG charts.
Staff-level multiple-choice quiz questions with instant feedback and answer explanations.
Firebase Google authentication automatically syncs your completed topics and quiz scores.
How clear and actionable was this distributed systems breakdown?