Home/Labs/Error Budget Policy Game
All 280 Labs
INTERACTIVE LAB💰

Error Budget Policy Lab (Interactive)

Spend a monthly error budget on a bad deploy and trigger the automated freeze. Compute budget from SLO and traffic, burn it with a buggy canary, watch the release-gate policy state machine, and let the rolling window heal.

Error Budget Policy & Feature Freeze Engine

Budget = (1 − SLO) × traffic. Spend it on a bad deploy, watch the CI/CD gate slam shut, then let the rolling window heal.

Remaining budget @ day 0100.0%
Full-window budget50,000 errors
Consumed so far0
Burn rate (vs 1.0x fair)n/a
Downtime equivalent0.0 min lost of 43.2 min
GREEN · >50% leftFull velocity: multiple deploys/day, chaos experiments allowed, fast-track PR approvals.
GOVERNANCE LOG:

› Fresh 30-day rolling window: full error budget available. Spend it on canaries, chaos experiments and fast merges.

How It Works Under the Hood

Google SRE resolved the product-versus-platform war by declaring 100% reliability the wrong target: the error budget equals 100% minus the SLO, and it exists to be spent on risky releases, migrations, and chaos experiments. At 50M requests and a 99.9% SLO, a canary leaking 10,000 500s spends 20% of the month. Policy follows mechanically from the balance — above 50% deploy freely, under 20% mandate canaries and stop chaos, at zero the CI/CD Prometheus gate blocks non-security releases and the team spends the freeze on the reliability work that burned the budget.

Core Architectural Principles

  • Budget errors = monthly requests x (1 - SLO); 99.9% of 20M is 20,000 failures, no debate required.
  • Policy states drive action: green fast-track, orange canary-only, red automated feature freeze.
  • A 30-day rolling window heals automatically as burned hours age out of the measurement.
Interview Round Script

Frame error budgets as the governance artifact that converts uptime arguments into shared math: product and engineering both sign the policy before anyone is tired at 3 AM. Explain the freeze protocol and the Prometheus CI gate, and mention the 50% toil cap — if a service generates more than half operational toil, the pager goes back to the feature team.

Key Trade-Offs

Budgeted risk accelerates innovation with a hard safety stop, but only works with trustworthy telemetry and leadership that refuses to override the freeze.

Related Curriculum Chapter

Error Budgets & Site Reliability Engineering (SRE) Principles

Read Full Chapter Blueprint

Explore More Interactive Labs

View All 280 Labs