When to Break a Monolith Lab (Interactive)
Score six readiness gates before extracting a service, then watch failure probability grow with fan-out. A checklist-driven lab: clear all gates to extract, miss them and the simulator argues for staying modular. Also computes 1-(1-p)^N cascade risk.
Monolith Extraction Readiness Gates
Tick the service-extraction checklist and stress-test the distributed-monolith failure math before cutting a boundary.
Premature fan-out damage model
A distributed monolith inherits this multiplication plus lockstep deployments and shared-database backdoors. Extract only when the six readiness gates convert this chain into independent, observable, database-private services.
How It Works Under the Hood
Breaking a monolith is easy; breaking it correctly requires evidence. A domain earns extraction when it has distinct scaling needs, its own data ownership, an independent release cadence, on-call clarity, and stable interface boundaries. Without those gates, you convert one process with in-process isolation into a distributed system where each extra dependency multiplies failure probability: P(fail) = 1-(1-p)^N. The math says a ten-hop chain at 0.5% per-hop failure fails once in twenty requests, which is a reliability tax paid before any velocity gain arrives.
Core Architectural Principles
- Six readiness gates (scaling, data, cadence, ownership, interface, team) must clear before extraction.
- Fan-out failure math P_fail = 1-(1-p)^N shows why more services means more frequent partial failures.
- Verdict logic: six gates ready means extract, four to five means hold, fewer means keep the monolith modular.
In system design interviews, when asked "should we split this service?", answer with gates, not ideology: distinct scaling profile, own database, independent deploy cadence, clear on-call. Then cite the fan-out failure math to show you understand the cost side, and propose the strangler fig rather than a rewrite if the gates pass.
Extraction buys team velocity only when readiness gates pass; otherwise you trade in-process reliability for distributed complexity.