Learning Rate
The gradient tells the model which way to walk, but not how far. That distance is the learning rate η — the single most sensitive knob in machine learning. Too small and training crawls; too large and it oscillates, then explodes to NaN. This topic covers why, what each failure looks like on a loss curve, the schedules (warmup + decay) that fix both, and cheap ways to find the right value.
01.The Problem: The Gradient Gives Direction, Not Distance
Remember the gradient update from Topic 11:
w ← w − η ∇J(w)
The gradient ∇J(w) is a compass needle: it says "this way is steepest downhill".
But a needle has no legs. It never says how far to step.
That distance is η (eta), the learning rate — one number you must choose yourself, and the single most sensitive hyperparameter in all of machine learning.
Why is it so sensitive? Because the gradient's advice is only valid locally.
∇J(w) was computed at your current position. The linear approximation of J is accurate only in a shrinking neighborhood around it. Walk two tiny steps and the right direction may have changed completely.
So η trades one thing against another:
- Big η → more progress per step, but you leave the region where the gradient meant anything → unreliable steps.
- Small η → reliable steps, but progress per step is nearly zero → thousands of wasted epochs.
Pick the extreme either way and you lose:
- Too small: the model looks "fine" but is quietly undertrained.
- Too large: the loss stalls, spikes, then becomes NaN and the run is dead.
So the question becomes
How far can I walk along a direction I only trust for a few steps?
That question has a real mathematical answer (curvature, below), a practical ritual (schedules and range tests), and a lot of expensive production folklore. This topic is all three.
Step Size Regimes and Schedules 🎚️
Step Size Regimes and Schedules 🎚️
Below the optimum, η costs wall-clock; above the stability bound it costs the run entirely. Schedules exist because the right η changes as training progresses.
Unlock Topic #51: Learning Rate
You are viewing a preview. The full in-depth technical walkthrough, worked derivations, and code notebooks for this concept, along with self-assessment quizzes, are available with Pro or Lifetime Access.
Failure modes, high-throughput bottlenecks, and real FAANG implementation decisions.
Interactive system topology diagrams, live parameter simulators, and downloadable SVG charts.
Staff-level multiple-choice quiz questions with instant feedback and answer explanations.
Firebase Google authentication automatically syncs your completed topics and quiz scores.
How clear and actionable was this distributed systems breakdown?