TOPIC #90Intermediate 12 min read

Learning-Rate Schedulers: Warmup, Cosine Decay, Step Decay

AI
AI & ML Editorial
Report an issue
Key takeawayCore Concept Summary

The learning rate is not a single number — it is a curve you draw over training. You start small (warmup), climb to a peak, then ease down (decay). This topic explains why that shape became the default, with the formulas, the schedules, and the real 2026 settings.

01.The Problem: One Step Size Cannot Fit the Whole Run

You already met the learning rate (lr). In one plain sentence: the learning rate is how big a step the model takes when it updates its weights each round.

Big step → learns fast, but can overshoot or blow up.

Small step → safe, but takes forever.

Now picture a real training run: 100,000 steps.

At step 1 the model is total garbage. Its weights are random. Gradients are huge and noisy.

At step 99,000 the model is almost done. It is nudging weights by tiny amounts to polish the result.

So the question becomes

Insight

Is one learning rate really right for both moments?

If you use a big lr early, the noisy first gradients can make the loss explode (topic 86) or just diverge — the run is dead.

If you use a small lr, the early phase crawls and you waste the run.

If you keep lr fixed to the end, you never settle cleanly into a sharp, well-generalizing minimum — you keep bouncing around the bottom.

The fix is obvious once you say it out loud:

Insight

Let the learning rate change over time.

That changing value is a schedule. The learning rate stops being a number and becomes a curve.

A Modern LR Curve

PRO Architecture Blueprint

A Modern LR Curve

Warmup protects the unstable first steps; the decay shape (cosine, step, linear) then trades exploration for convergence. LLM pretraining = warmup + long cosine.

A Modern LR Curve
100%
Touchpad: Pinch to zoom • Drag to pan
Rendering visual architecture flowchart...
PRO & LIFETIME CURRICULUM

Unlock Topic #90: Learning-Rate Schedulers: Warmup, Cosine Decay, Step Decay

You are viewing a preview. The full in-depth technical walkthrough, worked derivations, and code notebooks for this concept, along with self-assessment quizzes, are available with Pro or Lifetime Access.

Production Deep Dive

Failure modes, high-throughput bottlenecks, and real FAANG implementation decisions.

Interactive Blueprints

Interactive system topology diagrams, live parameter simulators, and downloadable SVG charts.

Knowledge Assessment

Staff-level multiple-choice quiz questions with instant feedback and answer explanations.

Cross-Device Progress Sync

Firebase Google authentication automatically syncs your completed topics and quiz scores.

Rate This Architecture ChapterFeedback & Rating

How clear and actionable was this distributed systems breakdown?

Related Concepts & Cross-References

Indexed from curriculum