Home/Labs/Alert Fatigue Shift Lab
All 280 Labs
INTERACTIVE LAB🚨

Multi-Window Burn Rate Alerting Lab (Interactive)

Live a 12-hour on-call shift: choose naive or burn-rate rules and count the pages. A deterministic shift with one transient blip, one real outage, and one pod cascade. Tune alert policy, grouping, and inhibition to control the pager.

On-Call Shift Simulator: Symptom vs Cause Alerting

One shift, one transient blip, one real 45-minute outage. Choose the alert policy and see who gets paged at 3 AM.

00:00 shift start06:0012:00 shift end
Pages this shift1
False / non-actionable0
Minutes to real page28.8
Error budget consumed100.0%
Alert fatigue index: 0% of pages were noiseMulti-window fast burn (14.4x over 5m AND 1h) clears the blip entirely and pages at ~28.8 min once the long window confirms sustained 30x burn — a 2% budget loss per hour of real outage.

Golden rule: never page a human unless a customer-facing symptom is broken AND a human action helps. Both the short and the long window must burn simultaneously, so a 4-minute spike that self-clears drops below threshold before the pager ever fires.

How It Works Under the Hood

Alert fatigue makes engineers acknowledge pages without reading them, so the real outage dies buried under CPU noise. Symptom-based alerting pages only when users are broken; Google SRE multi-window multi-burn-rate rules go further: a 14.4x fast-burn page requires both a 5-minute and a 1-hour window above threshold, which fires in under half an hour on a genuine outage yet instantly clears for a 4-minute blip. Alertmanager then collapses 200 per-pod alerts by grouping and inhibits the 350-alert downstream cascade when the root-cause network alert is already firing.

Core Architectural Principles

  • Fast burn: 14.4x rate over 5m AND 1h windows consumes 2% of a 30-day budget and pages 24/7.
  • Requiring both short and long windows lets a self-clearing spike drop below threshold before the page.
  • Grouping, inhibition, and silences turn a rack failure into one incident instead of 550 notifications.
Interview Round Script

State the paging rule first — never wake a human unless a customer symptom is broken and action exists — then derive multi-window math: at 99.9% SLO, 1.44% errors burn 2% of budget per hour, and the dual-window check kills false positives. Close with Alertmanager grouping and inhibition to show you have thought about the storm, not just the rule.

Key Trade-Offs

Naive thresholds page fast but train engineers to ignore the pager; multi-window burn rules are precise but add minutes of detection latency and PromQL complexity.

Related Curriculum Chapter

Monitoring & Alerting Design: Fighting Alert Fatigue

Read Full Chapter Blueprint

Explore More Interactive Labs

View All 280 Labs