A/B Testing & Shadow Deployments for Models
A better notebook score does not prove a better product. This topic is the evidence ladder models climb into production: shadow (dark) traffic as the zero-risk first gate, canary and A/B experiments with a sound unit of randomization and guardrail metrics, and the statistical traps — peeking, sample-ratio mismatch, offline-online gaps — that ship fake wins.
01.The Problem: Offline Win, Online Shrug
Your team trained a new ranking model. Offline AUC improved by 0.008. Everyone is pleased.
Then the question nobody can answer without help:
Will real users actually click, book, or watch more if we swap it in?
Here is the uncomfortable part: offline numbers routinely fail to translate. Your test set is a frozen snapshot with every label filled in. Real traffic is alive: users see results in different positions, react to latency, change their behavior because the model changed theirs, and bring in populations your snapshot never saw. A +0.008 AUC that moves zero clicks is the norm in ML experimentation, not an anomaly.
So how do you prove a model is better where it matters — in production, on real people, without gambling their experience?
You need instruments that answer three different questions, in increasing order of risk and decreasing order of doubt:
- "Would it break anything?" → shadow deployment
- "Is it safe at small real exposure?" → canary
- "Did it cause the metric to move?" → A/B experiment
That sequence is the evidence ladder. Each rung produces the permission slip for the next.
Evidence Ladder: Shadow -> Canary -> A/B
Evidence Ladder: Shadow -> Canary -> A/B
Shadow mirrors live traffic with zero exposure; canary takes real traffic at small share behind guardrails; the A/B experiment randomizes to attribute business impact — each rung is a gate for the next.
Unlock Topic #233: A/B Testing & Shadow Deployments for Models
You are viewing a preview. The full in-depth technical walkthrough, worked derivations, and code notebooks for this concept, along with self-assessment quizzes, are available with Pro or Lifetime Access.
Failure modes, high-throughput bottlenecks, and real FAANG implementation decisions.
Interactive system topology diagrams, live parameter simulators, and downloadable SVG charts.
Staff-level multiple-choice quiz questions with instant feedback and answer explanations.
Firebase Google authentication automatically syncs your completed topics and quiz scores.
How clear and actionable was this distributed systems breakdown?