Home/Labs/Partial Failure Rack
All 280 Labs
INTERACTIVE LAB🚢

Network Partitions and Partial Failure Handling Lab (Interactive)

One slow dependency, shared threads, lockstep retries: watch partial become total and stop it. Fan out requests to three dependencies, degrade one, and combine retry strategy, circuit breakers, and bulkhead pools to keep goodput and p99 intact.

Partial-Failure Containment Rack

One of three dependencies turns slow. Retry storms and shared thread pools turn its problem into everyone’s problem — backoff, circuit breakers, and bulkheads are how you keep a partial failure partial.

Dependency health
Retry strategy
Outbound calls/sec
330
Threads held / capacity
1076 / 170
Goodput
66%
p99 latency
13040 ms
  • Users DBpool 20serving
  • Inventorypool 100saturated
  • Recommendationspool 50serving

Bulkheads held: only Inventory degraded (~66% goodput — render the page without it). Add exponential backoff with jitter so retries decorrelate, and a breaker so callers stop queueing into a saturated pool.

At 100 req/s × 3 fan-out with 2 retries: 330 calls/s — retry amplification is multiplication, not addition.

How It Works Under the Hood

In a partitioned network the failure is partial — Recommendations times out while Users DB is fine — but naive systems totalize it: shared thread pools let one slow call drain capacity from every handler, and fixed-interval retries synchronize clients into retry storms that multiply load on the struggling dependency exactly when it can least absorb it. Resilience is containment: exponential backoff with jitter decorrelates retries, circuit breakers fast-fail so callers stop queueing into saturation, and bulkheads bound each dependency’s threads so one failure domain cannot starve others. This rack models all three knobs against thread occupancy and goodput.

Core Architectural Principles

  • Retry amplification computed as fan-out times attempts per request, degrading precisely under stress.
  • Shared versus per-dependency pools: collateral healthy-service failures counted explicitly.
  • Breaker open-state trades a dependency’s traffic for freed threads and bounded p99.
Interview Round Script

Design reviews should hear defaults: “retries only on idempotent calls, exponential backoff with full jitter, breaker after sustained failure ratio, per-dependency thread budgets, timeouts shorter than caller SLA.” Then address partition asymmetry: minority side sheds load rather than assuming the majority runs.

Key Trade-Offs

Every resilience knob sheds some traffic or freshness to ensure one bad dependency never downs everything.

Related Curriculum Chapter

Network Partitions & Partial Failure Handling

Read Full Chapter Blueprint

Explore More Interactive Labs

View All 280 Labs