TOPIC #84Advanced 11 min read

Weight Initialization: Xavier & He

AI
AI & ML Editorial
Report an issue
Key takeawayCore Concept Summary

Training outcomes are decided before the first gradient flows: the distribution you draw initial weights from controls whether signals survive 50 layers. Zero init never breaks symmetry; too-small init fades the signal; too-large init explodes it. Xavier and He compute the exact variance that keeps every layer's gain at 1.

01.The Problem: The Dice You Roll Before Training Decide Whether Training Works

You've built a 50-layer network. You've written the training loop. You run it, and the loss... just sits there. Or goes to NaN. Or the predictions are the same value for every input.

Nothing in your code is wrong. The weights were.

Here's the plain-sentence setup you need: at step 0, the network's weights are random numbers you sample from some distribution — and that choice happens before any gradient has ever flowed.

So the question becomes

Insight

What distribution should the starting weights come from?

"Random small numbers, whatever" is not an answer — and at depth it's fatal. Think about what a 50-layer forward pass actually is (Topic 82): a product of 50 matrices:

a^(L) ≈ (W_L·φ'·…·W_1)·x

Every layer multiplies the signal's scale by some gain. A gain of 0.9, fifty times: 0.9⁵⁰ ≈ 0.005 — the signal is gone. A gain of 1.1, fifty times: 1.1⁵⁰ ≈ 117 — the signal is a roaring saturation mess. Only a gain of ≈ 1.0 per layer lets the input survive to the output and lets gradients survive back to the first layer.

Whether the product vanishes or explodes is decided almost entirely by the variance of the distribution you sampled weights from. Two researchers worked out the arithmetic: Glorot & Bengio (2010) and He, Zhang & Ren (2015). Their rules — Xavier and He — are the difference between "deep nets were impossible" and "deep nets are a library call."

Initialization Failure Modes and the Two Fixes

PRO Architecture Blueprint

Initialization Failure Modes and the Two Fixes

Wrong init dies silently: symmetry collapse, vanishing forward signal, or explosion. Xavier and He compute the variance that keeps activations healthy at scale.

Initialization Failure Modes and the Two Fixes
100%
Touchpad: Pinch to zoom • Drag to pan
Rendering visual architecture flowchart...
PRO & LIFETIME CURRICULUM

Unlock Topic #84: Weight Initialization: Xavier & He

You are viewing a preview. The full in-depth technical walkthrough, worked derivations, and code notebooks for this concept, along with self-assessment quizzes, are available with Pro or Lifetime Access.

Production Deep Dive

Failure modes, high-throughput bottlenecks, and real FAANG implementation decisions.

Interactive Blueprints

Interactive system topology diagrams, live parameter simulators, and downloadable SVG charts.

Knowledge Assessment

Staff-level multiple-choice quiz questions with instant feedback and answer explanations.

Cross-Device Progress Sync

Firebase Google authentication automatically syncs your completed topics and quiz scores.

Rate This Architecture ChapterFeedback & Rating

How clear and actionable was this distributed systems breakdown?

Related Concepts & Cross-References

Indexed from curriculum