Retry & Backoff Jitter Lab (Interactive)
Crash a recovering server with synchronized client retries, then fix it with Full Jitter and watch the waves flatten. Compare immediate retries, no-jitter exponential backoff, and AWS Full Jitter across a client fleet: peak QPS per attempt wave, traffic multiplier, and recovery success all diverge.
Retry Storm vs Full Jitter Wave Physics
A 200ms blip knocks 10,000 clients offline. Choose a retry policy and watch the peak hammering QPS the recovering server absorbs.
Full Jitter spreads each retry uniformly over [0, min(MaxSleep, base × 2^attempt)], so 10k clients arrive as smooth background noise. Marc Brooker's AWS research: lowest server load AND fastest client completion. Add a ≤10% retry budget and idempotency keys before retrying any POST.
How It Works Under the Hood
Retries are load amplifiers: with three retries every failed request can become four, and undisciplined strategies synchronize the pain. Immediate retry hammers a server that just hiccuped at full client count. Exponential backoff spreads attempts across growing windows but no-jitter schedules produce thundering herds at each window boundary — everyone sleeps 2 seconds and wakes together. AWS's Full Jitter picks a random delay in 0-to-min(cap, base×2^attempt), provably decorrelating the fleet so recovery waves spread instead of stacking. Retry budgets and idempotency keys are the guardrails that make retries safe.
Core Architectural Principles
- Per-attempt wave peaks: clients divided by the sleep window; synchronized windows create enormous spikes.
- Full Jitter spreads retries uniformly across 0-min(cap, base×2^attempt), flattening each wave by design.
- Amplification math: three retries under failure double or quadruple offered load exactly when capacity dips.
Say three sentences that signal production experience: cap retries and use a retry budget like Envoy or gRPC, never retry non-idempotent operations without idempotency keys, and always jitter — Full Jitter beats equal jitter in AWS's own benchmarks. Also mention retry budgets live one layer up: clients retry at most 10-20% of original traffic.
Retries convert transient failures into client-side availability, but synchronized or unbounded retries manufacture outages.