TOPIC #113Intermediate 12 min read

Vanishing & Exploding Gradients in RNNs

AI
AI & ML Editorial
Report an issue
Key takeawayCore Concept Summary

To learn from a long sequence, the gradient must multiply one factor per time step. If each factor shrinks the signal, early tokens get no credit (vanishing); if each grows it, the update detonates (exploding). Follow the product with small numbers, then see which fixes actually worked.

The Gradient Multiplies One Jacobian Per Time Step ⚖️

Credit from a loss at step T back to step 1 is a product of T near-identical matrices. Products of numbers do not decay politely — they vanish or detonate exponentially.

The Gradient Multiplies One Jacobian Per Time Step ⚖️
100%
Touchpad: Pinch to zoom • Drag to pan
Rendering visual architecture flowchart...

01.The Problem: The Answer at Step 500 Depends on Step 1 — Can Blame Travel That Far?

Recall Topic 111: an RNN learns by backpropagating through time. If a wrong prediction at step 500 was partly caused by a word at step 1, the gradient — the "which knob, in which direction" signal from the loss — must flow back 499 steps to adjust the weights.

Gradient descent is a credit-assignment machine. The question this topic answers:

Insight

Does credit for a distant event survive the trip backward through hundreds of cells?

The punchline, up front: usually no. Not because of a bug in your code, but because of arithmetic that is baked into the chain rule. Once you see the one-line derivation below, you will understand why every sequence architecture after 1997 — LSTM, GRU, attention, Transformers — can be read as a fix for exactly this.

02.The Idea in Plain Words: The Gradient Is a Product, One Factor Per Step

To teach the network that "the word at step 1 determined the answer at step 500," BPTT must compute ∂L/∂h₁. The chain rule through the unrolled recurrence gives:

∂L/∂h₁ = ∂L/∂h_T · ∏_{t=2}^{T} ∂h_t/∂h_{t-1}

with each factor

∂h_t/∂h_{t-1} = diag(tanh′(z_t)) · W_h

That product sign ∏ means: multiply T−1 of these factors together. Two facts make this the defining pathology of vanilla RNNs:

  1. Every factor contains the same W_h — the recurrent matrix, reused at every step. That is weight sharing (Topic 111): the blessing becomes the curse. Roughly, the whole product scales like ‖diag(tanh′)·W_h‖^{T-1}.
  2. Every factor squashes: 0 < tanh′(z) ≤ 1, and in practice the states saturate, pushing derivatives well below 1.

Combine them into one number — the effective per-step gain:

γ = λ_max(W_h) · mean(tanh′)

  • If γ < 1, gradient magnitude decays as γ^T. With a typical γ ≈ 0.7, a 50-step dependency sees gradients ~10⁻⁸ of their starting size. Learning across it is numerically impossible.
  • If γ > 1, the same exponential runs the other way — the update magnitude detonates.

Products of numbers do not decay politely. They vanish or they detonate.

03.A Simple Worked Example: 0.9 vs 1.1, Walked Line by Line

Pretend the gradient is one number and each step back multiplies it by a constant gain. Start with a credit signal of size 1.0 at step T.

A pessimistic-but-typical gain of 0.9 per step:

  • 10 steps back: 0.9^10 ≈ 0.35 — already two-thirds gone.
  • 20 steps: 0.9^20 ≈ 0.12
  • 50 steps: 0.9^50 ≈ 0.005 — the signal is 200x weaker.
  • 100 steps: ≈ 0.000027 — for float32 training, effectively zero.

An over-eager gain of 1.1 per step:

  • 10 steps: 1.1^10 ≈ 2.6
  • 50 steps: 1.1^50 ≈ 117
  • 100 steps: ≈ 13,780 — one gradient now swamps every other weight update.

And with γ = 0.7 (realistic once tanh saturates): 0.7^50 ≈ 2e-8 — matches the textbook number above.

Insight

Notice what did and did not matter.

It was not the first factor being tiny. It was the same factor multiplied over and over. This is why the vanilla RNN’s practical memory stayed at 5-10 steps even though its hidden state technically "contains" the whole past: the learning signal, not the information, is what dies.

04.Visual Intuition: One Narrow Corridor, Two Ways to Fail

Every piece of credit for step 1 must squeeze through the same corridor of multiplied Jacobians:

code
 loss at step T
      │  ∂L/∂h_T = 1.0            (credit starts full-size)
      ▼
  × diag(tanh′)·W_h   gain γ
      ▼  0.76 → 0.53 → 0.37 ...   each arrow = one step back
      ▼
  × diag(tanh′)·W_h
      ▼
 credit at h₁:
   γ < 1:  1.0 → 0.3 → 0.05 → ~0     📉 VANISHED  (silent)
   γ > 1:  1.0 → 3 → 9 → 13,780       📈 EXPLODED  (loud: NaN)
   γ = 1:  credit survives   ← what LSTMs engineer the corridor to do

Read the two failure rows like a doctor:

  • Vanishing ↓: the corridor is a maze of soft walls; the message fades before step 1 hears it. The model trains happily on short patterns and is structurally blind to long ones.
  • Exploding ↑: the corridor is a feedback screech; a few steps push states into unsaturated regions where gains exceed 1, updates overshoot, loss spikes, oscillates, or becomes NaN.

Same product, two signs of the same inequality. That is the whole story of this topic in one picture.

05.The Analogy: The Whisper Game (Telephone)

Thirty kids sit in a line. Kid 30 whispers a sentence to kid 29, who whispers to 28, and so on to kid 1.

  • If every kid repeats 90% as loudly as they heard it — by kid 20 the whisper is inaudible. That is the vanishing gradient: the blame for a bad final answer cannot reach the early weights. The teacher looks at kid 30’s wrong sentence and cannot tell kid 3 what to change.
  • If every kid repeats 110% as loudly — the last exchanges are shouting, and one correction drowns out everything else. That is the exploding gradient.
  • The fix humans would pick: a loyal repeater in the line — a kid who whispers the old message at full volume, untouched, alongside the new word. That is exactly what the LSTM cell state is engineered to be (next topic): a corridor where gain = 1.
  • Or: skip the line. Kid 30 sends kid 3 a text message — one hop, full volume, no matter how many kids sit between. That is attention: dependency distance collapses to one edge.

And when one kid shouts, the teacher says "everyone, keep your voice at most this loud." That is gradient clipping.

06.Vanishing vs Exploding: Two Faces, Different Fixes

Vanishing (γ < 1): the silent failure. Loss trains fine on short patterns; the model simply never learns long-range structure. You diagnose it behaviorally: a copy-task accuracy cliff past ~5-10 steps; BLEU collapse on 30+ word translation segments — exactly what Bahdanau’s 2014 paper measured, motivating attention.

Exploding (γ > 1): the loud failure. Updates overshoot — loss spikes, oscillates, or becomes NaN.

Exploding gradients have a brutally effective patch introduced with the 2013 ICML best paper (Pascanu et al.): norm clipping — rescale the full gradient whenever ‖g‖ > τ (typical τ = 1-10), preserving direction while capping step size. It is trivial, theoretically justified (bounded update, monotone improvement per step), and still standard in every LSTM/GRU/Transformer training loop in 2026.

Vanishing gradients got no such cheap fix — they required new architectures.

07.Consequence: Credit Assignment Across Time — Two Different Questions

Even when information persists in the hidden state (Topic 112 argues it often does not), the learning signal to exploit it cannot arrive. This splits the failure into two distinct questions an interviewer may probe:

  1. Representation question: can h_t encode a fact from step 1? Sometimes.
  2. Optimization question: can SGD discover weights that use that encoding? Almost never at T > ~10 with vanilla recurrence — the gradient of every relevant parameter is a coin-flip of numerical noise by then.

The historical response order is instructive:

↓ first gating (Hochreiter & Schmidhuber 1997, popularized ~2014) created additive gradient highways; ↓ then attention (2014) short-circuited the path entirely by connecting distant positions with a single edge; ↓ residual connections (2015, adopted into Transformers 2017) made "skip one Jacobian" a structural default of every deep net.

python— Watch the decay: gradient magnitude through 40 tanh-RNN steps
import torch
torch.manual_seed(0)
d = 128
Wh = torch.randn(d, d) / d**0.5          # sane init: λ_max ≈ 1
g = torch.randn(d, 1)                     # dL/dh_T
mag = []
for t in range(40):                       # backprop h_T -> h_{T-40}
    dz = torch.randn(d, 1)                # typical pre-activation
    g = Wh.T @ (g * (1 - torch.tanh(dz)**2))   # diag(tanh') · W_h
    mag.append(g.norm().item())
print([f"{m:.1e}" for m in mag[::10]])   # e.g. ['3.1e-01', '1.8e-08', '6.0e-15']

08.Why AI Cares: The Full Mitigation Toolbox (What Actually Gets Used)

Ranked by how much of the exponential they neutralize:

  • Gradient clipping (exploding only): torch.nn.utils.clip_grad_norm_ — table stakes, keeps training from diverging; orthogonal to vanishing.
  • Architecture with additive paths — LSTM (next topic): the forget gate can hold ∂C_t/∂C_{t-1} ≈ 1, converting the product of decaying factors into a product of near-identity factors. Constant error carousel, explicitly designed as the vanishing-gradient fix (Gers, Schmidhuber & Cummins, 1999).
  • Attention/Transformers: dependency length becomes a single matmul hop regardless of distance — the product collapses to one Jacobian. Transformers did not make gradients big; they made the path short. Residual + LayerNorm keep it healthy (pre-LN in 2024-2026 defaults trains 100+ layers stably).
  • Initialization tricks: unit-norm / orthogonal recurrent init (Arjovsky et al., IRNN, 2016) can learn dependencies to ~hundreds of steps in idealized tasks — real text still favors gates or attention.

In practice, this pathology is why production stacks in 2026 clip gradients in every loop (the exploding half is never fully gone) and why "why did RNNs lose to Transformers?" has a gradient-shaped answer: whoever controls the product controls the memory.

Architectural Trade-offs & Production Realities

Architectural Advantages

  • Understanding this pathology explains every later design choice in the phase: gates, skip paths, attention, residuals.
  • Exploding side has a one-line fix (norm clipping) used unchanged since 2013.
  • The product-of-Jacobians lens also explains Transformer training needs (LayerNorm, warmup, √dk scaling).

Trade-offs & Constraints

  • Vanishing side cannot be fixed by hyperparameters alone in vanilla RNNs — architecture change is mandatory.
  • Truncated BPTT caps learnable range as a workaround, trading long dependencies for trainability.
  • Symptoms are quiet: models converge and post decent short-range metrics while being structurally blind to long context.
Production Implementation in Big Tech
Google Neural Machine Translation (pre-attention-era diagnosis)• Why RNN translation quality fell off with sentence length

Bahdanau et al. (2014) documented that fixed-vector seq2seq BLEU dropped sharply beyond ~30-word inputs — the behavioral signature of a gradient-starved, bottlenecked encoder. The diagnosis directly produced attention, and Google’s later production GNMT (2016) and Transformer (2017) both bake in the fix: never force long-distance credit through a T-step product.

Staff+ Engineering Takeaways

  • BPTT gradients through time are a product of T copies of diag(σ′)·W_h — magnitude scales exponentially with distance.
  • Per-step gain < 1 (typical, due to tanh saturation) → vanishing; > 1 → exploding (NaN/oscillating loss, fixed by norm clipping).
  • Vanishing is a *learning* failure distinct from the *representation* limit of the hidden state; both cap vanilla RNN memory at ~5-10 steps.
  • LSTM/GRU attack the problem with additive near-identity gradient paths; attention attacks it by making dependency paths one hop.
  • Clipping, good init, and truncation help at the margins — only architecture changes defeat the exponential.

Topic Knowledge Check

Exercise 1 of 3 • Test your architectural comprehension.

Exercise 1 of 30 answered
1

Why does weight sharing across time make vanishing gradients *worse*, not better?

Rate This Architecture ChapterFeedback & Rating

How clear and actionable was this distributed systems breakdown?