A/B Testing Infrastructure Lab (Interactive)
Hash real user cohorts into variants and price the sample size for significance. Implement MurmurHash3 bucketing live: assign thousands of users deterministically across experiment layers, then compute per-arm traffic and run time.
MurmurHash3 Bucketing & Power Analysis Lab
Assign a real user cohort with hash(userId + experimentId) mod 100 — no database — then price the sample size you need for a winner.
The avalanche effect in action: every user id is hashed with the experiment id mixed in, so assignment is deterministic (same user, same variant across devices and refreshes — zero DB lookups) yet uncorrelated across layers, which is what lets 3+ teams run concurrent orthogonal experiments like Booking.com\'s 1,000-test platform.
Peeking warning:N ∝ 1/MDE²: halving the detectable lift quadruples the required 451,316 users/arm. Stopping the test the first morning p < 0.05 flickers green inflates the false-positive rate from 5% to 30%+ — pre-commit the sample size, or switch to sequential testing. And keep guardrails on: a variant can win conversion while breaching the p99-latency or crash-rate floor, which auto-kills it back to control.
How It Works Under the Hood
A billion-user experimentation platform cannot ask a database who gets which checkout, so it derives assignment from math: bucket = MurmurHash3(userId + experimentId) mod 100 evaluates in under 50 nanoseconds, is perfectly deterministic across devices, and — because the experiment id is mixed into the key — produces orthogonal distributions per layer so hundreds of teams test concurrently without contamination. The statistical half is equally unforgiving: N per arm scales with 1/MDE squared, so halving detectable lift quadruples runtime, and peeking at p-values each morning inflates false positives from 5% toward 30%.
Core Architectural Principles
- Avalanche-hashed bucketing is deterministic, uniform, and storage-free: no user-to-variant table needed.
- Orthogonal layers re-hash with a layer key so the same user can be in independent concurrent tests.
- Sample size n = 15.68 x p(1-p) / delta^2 per arm at alpha 0.05 with 80% power; guardrails auto-kill breaches.
For experimentation questions, recite the assignment formula and its three wins: deterministic, zero-latency, multi-device consistent. Then show statistical literacy — pre-computed sample size, power, the peeking problem, and guardrail metrics like p99 latency and crash rate that veto a conversion win, because a variant that lifts purchases 2% while doubling latency is a loss.
Data-driven decisions beat opinions, but small-effect tests need millions of users and weeks of runtime, and unmanaged overlapping experiments leak bias.