Incident Management MTTR Lab (Interactive)
Command a live P1 outage: every decision burns minutes and user trust. Sequence war-room roles, mitigations, and mistakes against a running clock; the simulator scores MTTR, users impacted, and response quality.
P1 War Room: Command the Incident
Every click burns wall-clock minutes while ~1,200 users/min keep hitting 503s. Mitigate first — the MTTR math will judge you.
› 14:00 UTC · PagerDuty P1: checkout 5xx rate > 2%. Free-for-all in #general. Choose your first move.
How It Works Under the Hood
High-severity outages fail twice: first to the bug, then to uncoordinated debugging. The Incident Command System separates authority — the IC orchestrates and never touches a terminal, the technical lead investigates, the comms lead keeps Statuspage honest — around one cardinal rule: mitigate first, investigate later. A rollback in five minutes beats a heap dump in twenty-five. Hours later, the blameless postmortem runs the 5 Whys past the human who shipped v2.4 to the missing automated gate that let it through, because punished engineers hide the details you need next time.
Core Architectural Principles
- Role separation: IC holds authority and coordination, tech lead runs diagnosis, comms lead owns stakeholders.
- Mitigation order matters: feature-flag kill switch ~1 min, rollback ~5 min, live debugging 25+ min of outage.
- Blameless 5 Whys ends at systemic fixes with SMART action items, never at an individual.
When asked how you handle production incidents, narrate the protocol: declare, assign IC, page SME, mitigate with rollback or flag toggle, communicate every 15 minutes, then postmortem within 48 hours. Say the magic words — mitigate first, investigate later — and mention that unresolved postmortem action items should block feature shipping.
Structured command costs a few setup minutes and feels slow, versus free-for-all debugging that reliably multiplies MTTR and burns the wrong person out.