Dropout: Training an Ensemble by Deleting Neurons
Dropout randomly switches off a fraction of neurons at every training step, so no neuron can lean on any single partner. It works like training a huge ensemble that shares weights. This topic covers the mechanism, inverted scaling, where to place it, and its modern decline in LLM pretraining.
01.The Problem: Neurons Form Lazy Cliques
A network learns by tuning many neurons working together.
Here is a failure mode. Two neurons can start over-depending on each other. Neuron A only fires correctly because neuron B always says exactly the same thing. They form a clique.
That is fragile. On the training data the clique looks great. On new data, if neuron B behaves even a little differently, A's assumption breaks and the whole thing falls apart. This over-coupling is called co-adaptation, and it is a big cause of overfitting (memorizing the training set instead of learning general rules).
So the question becomes:
How do I force every neuron to be useful on its own, without any guaranteed partner?
The cheeky answer from Srivastava & Hinton (2014): randomly delete neurons while training. If a neuron never knows whether its partner will be there, it cannot build a clique. This is Dropout.
Dropout as Train-Time Ensembling
Dropout as Train-Time Ensembling
Each step trains a different thinned network (weight-shared); inverted dropout rescales at train time so inference is untouched. At test you evaluate the approximate ensemble mean.
Unlock Topic #93: Dropout: Training an Ensemble by Deleting Neurons
You are viewing a preview. The full in-depth technical walkthrough, worked derivations, and code notebooks for this concept, along with self-assessment quizzes, are available with Pro or Lifetime Access.
Failure modes, high-throughput bottlenecks, and real FAANG implementation decisions.
Interactive system topology diagrams, live parameter simulators, and downloadable SVG charts.
Staff-level multiple-choice quiz questions with instant feedback and answer explanations.
Firebase Google authentication automatically syncs your completed topics and quiz scores.
How clear and actionable was this distributed systems breakdown?