Skip and Residual Connections
Common sense said deeper networks are better — until deeper networks started getting worse. The fix was beautifully simple: let each layer ADD a small correction to its input instead of rewriting it from scratch (y = F(x) + x). That one wiring change cured the degradation problem, gave gradients a shortcut highway, and now sits inside every Transformer you have ever used.
01.The Problem: Deeper Got Worse
Build a 20-layer image network. It works well.
Now build a 56-layer version of the same recipe and — surely — better?
He et al. (2015) tried it and got the opposite: the 56-layer plain network had higher training error than the 20-layer one. Not test error. Training error — it was worse at the very images it was being trained on.
Pause on how strange that is.
A deeper net contains the shallower one. Couldn't the extra layers just... do nothing?
Exactly. If 20 layers are good, then "20 good layers + 36 no-op layers" should be at least as good. Doing nothing is the easiest job in the world. Yet gradient descent (the training algorithm that nudges every knob a little way downhill) could not find the "do nothing" answer inside a stack of ordinary layers. Each extra layer had to re-derive the whole signal from scratch, and the chain of weights became hard to steer. The optimizer got lost.
That failure mode has a name: the degradation problem. And its fix — an architectural wiring change, not a training trick — is the most copied idea in modern deep learning.
One residual block 🔗
One residual block 🔗
The block computes F(x) through two conv+norm layers, then adds the unmodified input x via the identity shortcut (orange). ReLU is applied after the addition. The residual F(x) = y - x is what the weights learn.
Unlock Topic #104: Skip and Residual Connections
You are viewing a preview. The full in-depth technical walkthrough, worked derivations, and code notebooks for this concept, along with self-assessment quizzes, are available with Pro or Lifetime Access.
Failure modes, high-throughput bottlenecks, and real FAANG implementation decisions.
Interactive system topology diagrams, live parameter simulators, and downloadable SVG charts.
Staff-level multiple-choice quiz questions with instant feedback and answer explanations.
Firebase Google authentication automatically syncs your completed topics and quiz scores.
How clear and actionable was this distributed systems breakdown?