Model Pruning
Most trained weights barely matter — pruning deletes them. But only certain deletion patterns convert into real speed: unstructured zeros need special kernels, 2:4 blocks ride NVIDIA hardware for ~2x, and structured cuts shrink the model outright. Here is the ladder, the one-shot LLM methods (SparseGPT, Wanda), and the honest math of when sparsity pays.
01.The Problem: Your Model Is Full of Noise
Here is a slightly offensive fact about trained neural networks.
Most of their parameters do almost nothing.
Look at the distribution of learned weight magnitudes: a big pile near zero, a thin tail of important values.
codecount ▄ █▄ ██▄ ███▄___ ─┴────┴────┴──► |weight| ↑ thousands of weights this close to zero that removing them changes the output barely at all
Back in 2015, Han et al. showed deep nets tolerate 70-90%+ of weights zeroed with retraining, still working. Later, the lottery-ticket hypothesis (2018) made it stranger: inside a randomly-initialized network there already exists a sparse subnetwork that trains to full accuracy.
So the question becomes
If we can delete most of the dead weight, why does everyone still ship the full model?
Two hard reasons, and this whole topic is about them:
- Deleting individual numbers does not automatically make math faster. A matrix with scattered zeros is still the same-sized matrix on a GPU (Section 4).
- LLMs are less redundant than old CNNs. What was safe at 80% sparsity in 2015 collapses a 70B model today past ~50-60%.
Pruning is the art of choosing which things to delete so that quality survives and — the honest part — speed actually materializes on your hardware.
Pruning Granularity Ladder ✂️
Pruning Granularity Ladder ✂️
From individual weights (high sparsity, needs kernels) to 2:4 blocks (hardware-native) to whole structural units (guaranteed dense speedup, quality risk) — granularity trades sparsity headroom against real-world acceleration.
Unlock Topic #183: Model Pruning
You are viewing a preview. The full in-depth technical walkthrough, worked derivations, and code notebooks for this concept, along with self-assessment quizzes, are available with Pro or Lifetime Access.
Failure modes, high-throughput bottlenecks, and real FAANG implementation decisions.
Interactive system topology diagrams, live parameter simulators, and downloadable SVG charts.
Staff-level multiple-choice quiz questions with instant feedback and answer explanations.
Firebase Google authentication automatically syncs your completed topics and quiz scores.
How clear and actionable was this distributed systems breakdown?