Bulkhead Isolation Lab (Interactive)
One thread pool or three compartments? A slow fraud dependency floods whichever design you give it. Allocate worker threads across Orders, Catalog, and a slow Fraud service using Little's Law, then watch a shared pool saturate while partitioned pools stay healthy.
Bulkhead Thread-Pool Compartments
Size pools with Little's Law (L = λ × W), then make the external fraud API hang and see who sinks.
Watertight compartments hold: even with the fraud API hanging, at most its 15% slice of threads blocks. Checkout and catalog keep serving — the ship stays buoyant.
How It Works Under the Hood
Ships survive sinking because watertight bulkheads confine flooding to one compartment; the software version partitions thread pools, connection pools, and queues per dependency. Apply Little's Law — threads needed equals arrival rate times service time — and the math is brutal: a fraud call at 2 seconds holding 20% of 4,000 QPS demands 1,600 threads. In a shared pool the slowest dependency drains every worker and even healthy endpoints queue behind it, so a third-party hiccup becomes a site-wide outage. Compartments size per dependency and cap the blast radius.
Core Architectural Principles
- Little's Law sizing: required threads per workload = its arrival share × QPS × its latency in seconds.
- Shared mode allocates from one pool, so high-latency Fraud work starves fast healthy endpoints.
- Partitioned mode splits the pool 60/25/15 per workload, flooding only the compartment that is slow.
When discussing resilience, sequence it: timeouts bound duration, circuit breakers bound attempt rate, bulkheads bound blast radius — you need all three. Make bulkheads concrete: a dedicated 50-thread pool and 1-second timeout for the payment gateway so it cannot poison your 200-thread core pool. Mention semaphores versus pools: semaphores cap concurrency without thread overhead.
Compartmentalized capacity isolates failure domains but fragments tuning and wastes idle threads per partition.