Multi-Window Burn Rate Alerting Lab (Interactive)
Live a 12-hour on-call shift: choose naive or burn-rate rules and count the pages. A deterministic shift with one transient blip, one real outage, and one pod cascade. Tune alert policy, grouping, and inhibition to control the pager.
On-Call Shift Simulator: Symptom vs Cause Alerting
One shift, one transient blip, one real 45-minute outage. Choose the alert policy and see who gets paged at 3 AM.
Golden rule: never page a human unless a customer-facing symptom is broken AND a human action helps. Both the short and the long window must burn simultaneously, so a 4-minute spike that self-clears drops below threshold before the pager ever fires.
How It Works Under the Hood
Alert fatigue makes engineers acknowledge pages without reading them, so the real outage dies buried under CPU noise. Symptom-based alerting pages only when users are broken; Google SRE multi-window multi-burn-rate rules go further: a 14.4x fast-burn page requires both a 5-minute and a 1-hour window above threshold, which fires in under half an hour on a genuine outage yet instantly clears for a 4-minute blip. Alertmanager then collapses 200 per-pod alerts by grouping and inhibits the 350-alert downstream cascade when the root-cause network alert is already firing.
Core Architectural Principles
- Fast burn: 14.4x rate over 5m AND 1h windows consumes 2% of a 30-day budget and pages 24/7.
- Requiring both short and long windows lets a self-clearing spike drop below threshold before the page.
- Grouping, inhibition, and silences turn a rack failure into one incident instead of 550 notifications.
State the paging rule first — never wake a human unless a customer symptom is broken and action exists — then derive multi-window math: at 99.9% SLO, 1.44% errors burn 2% of budget per hour, and the dual-window check kills false positives. Close with Alertmanager grouping and inhibition to show you have thought about the storm, not just the rule.
Naive thresholds page fast but train engineers to ignore the pager; multi-window burn rules are precise but add minutes of detection latency and PromQL complexity.