Home/Labs/Nines & Buffer Calculator
All 280 Labs
INTERACTIVE LAB🎯

SLA SLO SLI Nines Lab (Interactive)

Set the internal target and legal contract, then run an outage through the buffer. Convert availability percentages into allowed downtime per day, month, and year, and see exactly when an SLO breach still protects the SLA.

Nines Math: SLI → SLO → SLA Safety Buffer

Set the internal target and the legal contract, then drive an outage through the compliance window and watch which one breaks first.

SLO allowed month downtime4.3 min
SLA allowed month downtime7.20 h
Safety buffer7.13 h
Error budget (= 100 − SLO)0.010000000000005116% → 52.6 min/yr
SLO breached, SLA intact: the 30-min outage crossed the internal 4.3 min target, so engineers get paged and stabilize — with 6.70 h still to spare before the money breaks.
AvailabilityDown/monthDown/yearArchitecture required
99%7.20 h3.65 daysSingle server + automated backup restart
99.5%3.60 h1.83 daysMulti-instance behind a load balancer
99.9%43.2 min8.76 hLB fleet + read replicas (three nines)
99.95%21.6 min4.38 hMulti-AZ with automated failover drills
99.99% ← SLO4.3 min52.6 minMulti-AZ active-active, instant failover (four nines)
99.999%0.43 min5.3 minMulti-region active-active + Raft/Paxos (five nines)

Each extra nine multiplies infrastructure cost ~5-10x: four nines allows 4.38 min/month, five nines only 26.3 s — a full-region failover budget with no human in the loop. Measure over 30-day rolling windows, never calendar months, so a day-2 outage does not freeze releases for 29 days.

How It Works Under the Hood

The reliability triad is a hierarchy of accountability: the SLI is the measured ratio of good events to total valid events, the SLO is the internal engineering target over a 30-day rolling window, and the SLA is the contract whose breach pays service credits. Because each nine buys roughly 10x less downtime but costs 5-10x more infrastructure, the professional move is setting the SLO strictly tighter than the SLA: at 99.95% internal versus 99.0% contractual, a 30-minute outage pages engineers while customer refunds stay legally untouched, leaving a ~416-minute monthly mitigation runway.

Core Architectural Principles

  • Allowed downtime = (1 - availability) x window minutes: 99.9% is 43.8 min/month, 99.999% is 26.3 seconds.
  • The SLO-SLA gap is a deliberate safety buffer absorbing detection lag and synthetic-monitor jitter.
  • 30-day rolling windows beat calendar months: no 29-day release freeze after one early outage.
Interview Round Script

Differentiate the triad in one sentence each, then quote the downtime math cold — three nines is 8.76 hours a year, four is 52.56 minutes. Always add the buffer rule: internal SLO strictly tighter than contractual SLA, measured on rolling windows, because equal targets mean customers learn about outages before your alerts do.

Key Trade-Offs

Aggressive nines win enterprise contracts but multiply redundancy cost and legal liability; loose ones keep engineers unemployed but churn customers.

Related Curriculum Chapter

SLAs, SLOs, and SLIs: Measuring Service Reliability

Read Full Chapter Blueprint

Explore More Interactive Labs

View All 280 Labs