Home/Labs/Netflix Resilience Edge
All 280 Labs
INTERACTIVE LAB🍿

Netflix Resilience Edge Lab (Interactive)

Trip Hystrix breakers on a failing dependency and size the Open Connect edge plane. Drive error rates into a recommendation service wrapped in a 50-thread bulkhead, watch the circuit open and fall back to cached data, and split peak video traffic between ISP-embedded appliances and AWS origin.

Netflix Resilience & Edge Plane Lab

Trip Hystrix-style circuit breakers against a failing recommendation service, then size the Open Connect data plane that keeps 95% of video bytes out of AWS.

Breaker state: OPEN
Avg call latency1.0 ms
Threads busy (of 50-thread bulkhead)40+
Requests served cached fallback100%
Playback sessions degraded0%
Peak video bytes (data plane)87.5 Tbps
Bytes crossing back into AWS origin4375 Gbps (5% of video)
Circuit OPEN: recommendations serve cached fallbacks; playback threads stay free.

How It Works Under the Hood

Netflix runs hundreds of microservices behind the Zuul edge gateway with Eureka discovery, where a single hung dependency can exhaust a Playback service's thread pool and cascade into a full outage. Hystrix wraps each dependency in its own thread bulkhead and a three-state circuit breaker: CLOSED serves traffic, OPEN fast-fails to cached fallbacks, HALF-OPEN probes recovery. Meanwhile Open Connect appliances inside ISP networks deliver over 95% of video bytes, so the control plane and the data plane scale on completely different physics.

Core Architectural Principles

  • Circuit breaker state machine: error-rate thresholds flip OPEN -> HALF-OPEN -> CLOSED while cached Top 10 fallbacks keep the row playing.
  • Thread bulkhead math: concurrent threads consumed equals request rate times dependency latency, capped at 50 per dependency pool.
  • Data-plane bifurcation: 3.5 Mbps peak streams per subscriber terminate at Open Connect inside the ISP, leaving only 5% of bytes to AWS.
Interview Round Script

Frame Netflix as two planes: a latency-tolerant control plane protected by bulkheads, breakers, and graceful degradation, and an exabyte data plane pushed to the ISP edge. Quantify the cascade you prevent: at 40,000 QPS a 5-second timeout consumes the entire 50-thread pool in a fraction of a second, which is exactly why fallbacks beat retries.

Key Trade-Offs

Per-dependency isolation and cached degradation buy five-nines playback availability at the cost of thread-pool overhead and stale-but-present recommendations.

Related Curriculum Chapter

Netflix: Cloud-Native Microservices, Zuul, Hystrix, & Chaos Engineering

Read Full Chapter Blueprint

Explore More Interactive Labs

View All 280 Labs