Weight Initialization: Xavier & He
Training outcomes are decided before the first gradient flows: the distribution you draw initial weights from controls whether signals survive 50 layers. Zero init never breaks symmetry; too-small init fades the signal; too-large init explodes it. Xavier and He compute the exact variance that keeps every layer's gain at 1.
01.The Problem: The Dice You Roll Before Training Decide Whether Training Works
You've built a 50-layer network. You've written the training loop. You run it, and the loss... just sits there. Or goes to NaN. Or the predictions are the same value for every input.
Nothing in your code is wrong. The weights were.
Here's the plain-sentence setup you need: at step 0, the network's weights are random numbers you sample from some distribution — and that choice happens before any gradient has ever flowed.
So the question becomes
What distribution should the starting weights come from?
"Random small numbers, whatever" is not an answer — and at depth it's fatal. Think about what a 50-layer forward pass actually is (Topic 82): a product of 50 matrices:
a^(L) ≈ (W_L·φ'·…·W_1)·x
Every layer multiplies the signal's scale by some gain. A gain of 0.9, fifty times: 0.9⁵⁰ ≈ 0.005 — the signal is gone. A gain of 1.1, fifty times: 1.1⁵⁰ ≈ 117 — the signal is a roaring saturation mess. Only a gain of ≈ 1.0 per layer lets the input survive to the output and lets gradients survive back to the first layer.
Whether the product vanishes or explodes is decided almost entirely by the variance of the distribution you sampled weights from. Two researchers worked out the arithmetic: Glorot & Bengio (2010) and He, Zhang & Ren (2015). Their rules — Xavier and He — are the difference between "deep nets were impossible" and "deep nets are a library call."
Initialization Failure Modes and the Two Fixes
Initialization Failure Modes and the Two Fixes
Wrong init dies silently: symmetry collapse, vanishing forward signal, or explosion. Xavier and He compute the variance that keeps activations healthy at scale.
Unlock Topic #84: Weight Initialization: Xavier & He
You are viewing a preview. The full in-depth technical walkthrough, worked derivations, and code notebooks for this concept, along with self-assessment quizzes, are available with Pro or Lifetime Access.
Failure modes, high-throughput bottlenecks, and real FAANG implementation decisions.
Interactive system topology diagrams, live parameter simulators, and downloadable SVG charts.
Staff-level multiple-choice quiz questions with instant feedback and answer explanations.
Firebase Google authentication automatically syncs your completed topics and quiz scores.
How clear and actionable was this distributed systems breakdown?