SLA SLO SLI Nines Lab (Interactive)
Set the internal target and legal contract, then run an outage through the buffer. Convert availability percentages into allowed downtime per day, month, and year, and see exactly when an SLO breach still protects the SLA.
Nines Math: SLI → SLO → SLA Safety Buffer
Set the internal target and the legal contract, then drive an outage through the compliance window and watch which one breaks first.
| Availability | Down/month | Down/year | Architecture required |
|---|---|---|---|
| 99% | 7.20 h | 3.65 days | Single server + automated backup restart |
| 99.5% | 3.60 h | 1.83 days | Multi-instance behind a load balancer |
| 99.9% | 43.2 min | 8.76 h | LB fleet + read replicas (three nines) |
| 99.95% | 21.6 min | 4.38 h | Multi-AZ with automated failover drills |
| 99.99% ← SLO | 4.3 min | 52.6 min | Multi-AZ active-active, instant failover (four nines) |
| 99.999% | 0.43 min | 5.3 min | Multi-region active-active + Raft/Paxos (five nines) |
Each extra nine multiplies infrastructure cost ~5-10x: four nines allows 4.38 min/month, five nines only 26.3 s — a full-region failover budget with no human in the loop. Measure over 30-day rolling windows, never calendar months, so a day-2 outage does not freeze releases for 29 days.
How It Works Under the Hood
The reliability triad is a hierarchy of accountability: the SLI is the measured ratio of good events to total valid events, the SLO is the internal engineering target over a 30-day rolling window, and the SLA is the contract whose breach pays service credits. Because each nine buys roughly 10x less downtime but costs 5-10x more infrastructure, the professional move is setting the SLO strictly tighter than the SLA: at 99.95% internal versus 99.0% contractual, a 30-minute outage pages engineers while customer refunds stay legally untouched, leaving a ~416-minute monthly mitigation runway.
Core Architectural Principles
- Allowed downtime = (1 - availability) x window minutes: 99.9% is 43.8 min/month, 99.999% is 26.3 seconds.
- The SLO-SLA gap is a deliberate safety buffer absorbing detection lag and synthetic-monitor jitter.
- 30-day rolling windows beat calendar months: no 29-day release freeze after one early outage.
Differentiate the triad in one sentence each, then quote the downtime math cold — three nines is 8.76 hours a year, four is 52.56 minutes. Always add the buffer rule: internal SLO strictly tighter than contractual SLA, measured on rolling windows, because equal targets mean customers learn about outages before your alerts do.
Aggressive nines win enterprise contracts but multiply redundancy cost and legal liability; loose ones keep engineers unemployed but churn customers.