Home/Labs/Chaos Blast Radius Lab
All 280 Labs
INTERACTIVE LAB⚡

Chaos Engineering Blast Radius Lab (Interactive)

Pick a fault, size the blast radius, and see if the Big Red Button saves you. Run hypothesis-driven experiments against a simulated fleet: inject pod kills, latency, or packet loss, with a Prometheus-linked automatic abort guardrail.

Game Day: Fault Injection & Big Red Button

Hypothesis: killing part of the fleet will NOT move the 5xx needle. Set the blast radius, then try to disprove it safely.

Terminate random instances in the checkout Deployment
Projected peak 5xx0.55%
Users in blast radius10%
Error budget burn1.91%
Experiment length300s full run
Hypothesis holds: steady state (checkout 5xx < 1.0%) survived 10% pod killer because the replica buffer happened to cover the loss — push the radius further to find the breaking point.
Experiment ledger (business hours only)

No experiments yet. Netflix\'s Chaos Monkey lesson: run small, monitored, announced — and let the abort trigger, not a customer, end the test.

How It Works Under the Hood

Chaos engineering is the scientific method applied to reliability, not vandalism: define a steady-state metric, hypothesize that a specific failure will not disturb it, inject the fault in a controlled blast radius, and try to disprove the hypothesis while the whole team is online. Netflix walked that ladder from Chaos Monkey killing instances, through Latency Monkey and Chaos Gorilla taking down an AZ, to Chaos Kong evacuating a region in minutes. The safety invariant is the automated abort: if 5xx crosses the guardrail, eBPF filters are torn down and instances restored in milliseconds.

Core Architectural Principles

  • Steady-state baseline plus explicit hypothesis replaces guesswork with falsifiable experiments.
  • Fault classes: instance kill, injected latency via tc/eBPF, packet loss, resource stress, DNS and clock skew.
  • Blast radius starts at 1% of fleet; the Big Red Button aborts automatically on SLO guardrail breach.
Interview Round Script

Define chaos engineering as controlled hypothesis testing that finds the missing timeout before 3 AM finds it for you. Walk the Simian Army escalation — monkey to gorilla to kong — and stress the prerequisites: real SLO metrics, auto-scaling, and an abort trigger wired to Prometheus, or you are just causing outages with extra steps.

Key Trade-Offs

Proactive failure discovery builds deep resilience confidence, but experiments without monitoring-linked aborts and blast-radius limits manufacture the very outages they study.

Related Curriculum Chapter

Chaos Engineering: Netflix Chaos Monkey & Fault Injection

Read Full Chapter Blueprint

Explore More Interactive Labs

View All 280 Labs