Home/Labs/Pod Shutdown Race Lab
All 280 Labs
INTERACTIVE LAB🛑

Graceful Shutdown Race Lab (Interactive)

Tune preStop, drain time, and grace period; count the 502s you leak. The kubelet removes endpoints and delivers SIGTERM in parallel. Configure the lifecycle and the Little's-law math tallies refused and killed requests.

Pod Termination Race: preStop vs iptables Propagation

Endpoint removal and SIGTERM run in parallel. Configure the lifecycle and count the HTTP 502s you pay for getting it wrong.

t=0 Terminating+5s SIGTERM60s
Refused (502 storm window 0s)0 req
SIGKILLed mid-request0 req
Need preStop+drain13s vs grace 45s
Active in-flight (Little L=λW)80,000 concurrent-sec
Zero-downtime exit: the 5s preStop sleep fully covers the 3s propagation, the drain finishes inside the 45s grace period, and no in-flight work was abandoned. Kafka unsubscribed, pools closed, exit code 0.

How It Works Under the Hood

Kubernetes pod termination is a race by construction: the endpoint controller propagates IP removal through every node's iptables over 1-3 seconds while the kubelet simultaneously runs preStop and sends SIGTERM. A handler that closes the listening socket on SIGTERM starts refusing traffic that stale load balancers still route — pure 502s. The fix is arithmetic, not magic: a preStop sleep covering propagation, then connection draining with Connection: close inside terminationGracePeriodSeconds, because anything unfinished at the deadline meets the uncatchable SIGKILL.

Core Architectural Principles

  • SIGTERM (15) is catchable and starts draining; SIGKILL (9) is uncatchable and abandons in-flight work.
  • preStop sleep of ~5s covers the 1-3s EndpointSlice/iptables propagation window.
  • In-flight requests follow Little's law (L = λW); grace must exceed preStop plus the longest request.
Interview Round Script

For zero-downtime deployment questions, name the race explicitly: SIGTERM and endpoint removal are parallel and unordered, so you need a preStop sleep before the handler closes sockets. Then list the drain checklist — stop accepting new connections, Connection: close, finish in-flight, close DB pools, unsubscribe Kafka — and note maxUnavailable: 0 keeps capacity whole.

Key Trade-Offs

Graceful draining adds 10-30 seconds per pod and needs maxSurge headroom, versus the 502 storms, orphaned locks, and rebalance storms of abrupt kills.

Related Curriculum Chapter

Graceful Shutdown & Rolling Restarts

Read Full Chapter Blueprint

Explore More Interactive Labs

View All 280 Labs