Chaos Engineering Blast Radius Lab (Interactive)
Pick a fault, size the blast radius, and see if the Big Red Button saves you. Run hypothesis-driven experiments against a simulated fleet: inject pod kills, latency, or packet loss, with a Prometheus-linked automatic abort guardrail.
Game Day: Fault Injection & Big Red Button
Hypothesis: killing part of the fleet will NOT move the 5xx needle. Set the blast radius, then try to disprove it safely.
No experiments yet. Netflix\'s Chaos Monkey lesson: run small, monitored, announced — and let the abort trigger, not a customer, end the test.
How It Works Under the Hood
Chaos engineering is the scientific method applied to reliability, not vandalism: define a steady-state metric, hypothesize that a specific failure will not disturb it, inject the fault in a controlled blast radius, and try to disprove the hypothesis while the whole team is online. Netflix walked that ladder from Chaos Monkey killing instances, through Latency Monkey and Chaos Gorilla taking down an AZ, to Chaos Kong evacuating a region in minutes. The safety invariant is the automated abort: if 5xx crosses the guardrail, eBPF filters are torn down and instances restored in milliseconds.
Core Architectural Principles
- Steady-state baseline plus explicit hypothesis replaces guesswork with falsifiable experiments.
- Fault classes: instance kill, injected latency via tc/eBPF, packet loss, resource stress, DNS and clock skew.
- Blast radius starts at 1% of fleet; the Big Red Button aborts automatically on SLO guardrail breach.
Define chaos engineering as controlled hypothesis testing that finds the missing timeout before 3 AM finds it for you. Walk the Simian Army escalation — monkey to gorilla to kong — and stress the prerequisites: real SLO metrics, auto-scaling, and an abort trigger wired to Prometheus, or you are just causing outages with extra steps.
Proactive failure discovery builds deep resilience confidence, but experiments without monitoring-linked aborts and blast-radius limits manufacture the very outages they study.