Network Partitions and Partial Failure Handling Lab (Interactive)
One slow dependency, shared threads, lockstep retries: watch partial become total and stop it. Fan out requests to three dependencies, degrade one, and combine retry strategy, circuit breakers, and bulkhead pools to keep goodput and p99 intact.
Partial-Failure Containment Rack
One of three dependencies turns slow. Retry storms and shared thread pools turn its problem into everyone’s problem — backoff, circuit breakers, and bulkheads are how you keep a partial failure partial.
- Outbound calls/sec
- 330
- Threads held / capacity
- 1076 / 170
- Goodput
- 66%
- p99 latency
- 13040 ms
- Users DBpool 20serving
- Inventorypool 100saturated
- Recommendationspool 50serving
Bulkheads held: only Inventory degraded (~66% goodput — render the page without it). Add exponential backoff with jitter so retries decorrelate, and a breaker so callers stop queueing into a saturated pool.
At 100 req/s × 3 fan-out with 2 retries: 330 calls/s — retry amplification is multiplication, not addition.
How It Works Under the Hood
In a partitioned network the failure is partial — Recommendations times out while Users DB is fine — but naive systems totalize it: shared thread pools let one slow call drain capacity from every handler, and fixed-interval retries synchronize clients into retry storms that multiply load on the struggling dependency exactly when it can least absorb it. Resilience is containment: exponential backoff with jitter decorrelates retries, circuit breakers fast-fail so callers stop queueing into saturation, and bulkheads bound each dependency’s threads so one failure domain cannot starve others. This rack models all three knobs against thread occupancy and goodput.
Core Architectural Principles
- Retry amplification computed as fan-out times attempts per request, degrading precisely under stress.
- Shared versus per-dependency pools: collateral healthy-service failures counted explicitly.
- Breaker open-state trades a dependency’s traffic for freed threads and bounded p99.
Design reviews should hear defaults: “retries only on idempotent calls, exponential backoff with full jitter, breaker after sustained failure ratio, per-dependency thread budgets, timeouts shorter than caller SLA.” Then address partition asymmetry: minority side sheds load rather than assuming the majority runs.
Every resilience knob sheds some traffic or freshness to ensure one bad dependency never downs everything.