Error Budget Policy Lab (Interactive)
Spend a monthly error budget on a bad deploy and trigger the automated freeze. Compute budget from SLO and traffic, burn it with a buggy canary, watch the release-gate policy state machine, and let the rolling window heal.
Error Budget Policy & Feature Freeze Engine
Budget = (1 − SLO) × traffic. Spend it on a bad deploy, watch the CI/CD gate slam shut, then let the rolling window heal.
› Fresh 30-day rolling window: full error budget available. Spend it on canaries, chaos experiments and fast merges.
How It Works Under the Hood
Google SRE resolved the product-versus-platform war by declaring 100% reliability the wrong target: the error budget equals 100% minus the SLO, and it exists to be spent on risky releases, migrations, and chaos experiments. At 50M requests and a 99.9% SLO, a canary leaking 10,000 500s spends 20% of the month. Policy follows mechanically from the balance — above 50% deploy freely, under 20% mandate canaries and stop chaos, at zero the CI/CD Prometheus gate blocks non-security releases and the team spends the freeze on the reliability work that burned the budget.
Core Architectural Principles
- Budget errors = monthly requests x (1 - SLO); 99.9% of 20M is 20,000 failures, no debate required.
- Policy states drive action: green fast-track, orange canary-only, red automated feature freeze.
- A 30-day rolling window heals automatically as burned hours age out of the measurement.
Frame error budgets as the governance artifact that converts uptime arguments into shared math: product and engineering both sign the policy before anyone is tired at 3 AM. Explain the freeze protocol and the Prometheus CI gate, and mention the 50% toil cap — if a service generates more than half operational toil, the pager goes back to the feature team.
Budgeted risk accelerates innovation with a hard safety stop, but only works with trustworthy telemetry and leadership that refuses to override the freeze.