Knowledge Distillation
Compress capability, not just bytes: a big teacher model trains a small student by sharing its soft probability distributions (via temperature) or, in the LLM era, its written output as synthetic data. The Phi playbook made this the cheapest way to raise what a small model can do.
01.The Problem: The Good Model Is Too Expensive to Keep
You have a model that works beautifully. It also costs a small fortune per query: 70B parameters, eight GPUs, hundreds of milliseconds per response.
The business says: serve it to a million users, or put it on a phone.
So you need a smaller model. Options from earlier in this phase:
- Quantize it (INT4)? That stores the same model cheaper — quality has a floor.
- Prune it? Deleting structure risks silently destroying capabilities.
But there is a third road, and it is the oldest trick in education:
Hire a tutor.
Take the big model (the teacher), and train a small model (the student) by having the student copy the teacher's behavior — not just the textbook answers.
Hinton, Vinyals & Dean (2015) reframed model compression as exactly this: distribution matching. The student does not need the teacher's weights; it needs the teacher's opinions, and opinions carry far more information than right/wrong labels.
What you get at the end: a plain, small, dense model — 10-20x cheaper to serve — with most of the teacher's behavior, and zero runtime dependence on the teacher.
Teacher-Student Distillation Loop 🎓
Teacher-Student Distillation Loop 🎓
The student trains on the same inputs, matching the teacher's softened output distribution (dark knowledge) alongside or instead of hard labels; gradients never flow into the frozen teacher.
Unlock Topic #182: Knowledge Distillation
You are viewing a preview. The full in-depth technical walkthrough, worked derivations, and code notebooks for this concept, along with self-assessment quizzes, are available with Pro or Lifetime Access.
Failure modes, high-throughput bottlenecks, and real FAANG implementation decisions.
Interactive system topology diagrams, live parameter simulators, and downloadable SVG charts.
Staff-level multiple-choice quiz questions with instant feedback and answer explanations.
Firebase Google authentication automatically syncs your completed topics and quiz scores.
How clear and actionable was this distributed systems breakdown?