Home/Labs/Timeout & Deadline Lab
All 280 Labs
INTERACTIVE LAB⏳

Timeout Design Lab (Interactive)

Run a three-hop chain under one budget and see static timeouts burn zombie compute while deadlines fast-fail. Compare per-hop static timeouts against gRPC-style remaining-budget propagation across Gateway-A-B-C, and separate connection timeouts from read timeouts in a TCP blackhole drill.

Timeout & Deadline Propagation Lab

Gateway → A → B → C. Compare static per-hop timeouts against gRPC-style remaining-budget propagation.

Call-chain outcome

Service A · order API (gateway → A)800 msOK
Service B · inventory (A → B)1000 msOK
Service C · warehouse DB query (B → C)700 msDISCARDED

User sees

2000 ms

Zombie compute

700 ms

Saved by mode

0 ms

Chain needs

2500 ms

User got a 504 at 2000 ms, yet later hops kept grinding because no hop knows the parent expired: 700 ms of zombie CPU (DISCARDED rows) paid for responses nobody will read. This is exactly the wasted-second-hop example the topic calls out — fix it with grpc-timeout / X-Request-Deadline headers.

Connection-timeout layer: TCP blackhole drill

Handshake completes normally, so this worker is held ~1000 ms by the read timeout — connectTimeout is invisible on the happy path. Flip the host to unroutable to see why internal calls want 200–500 ms connect timeouts while read timeouts stay at p99.9 + 3σ (typically 1–3 s).

How It Works Under the Hood

Client libraries ship infinite timeouts by default, and one frozen dependency then parks every worker thread until the fleet dies. Three layers must be configured deliberately: a 200-500 ms connection timeout for TCP/TLS establishment, a read timeout at roughly p99.9 plus variance for server processing, and a total call budget. Static per-hop timeouts still waste compute — a gateway that already returned 504 at 2 seconds leaves downstream services grinding another full second for an abandoned response. Deadline propagation passes the remaining budget in headers so each service aborts instantly when the parent has expired.

Core Architectural Principles

  • Chain model computes per-hop abort state: OK, ABORTED at timeout, or DISCARDED after the user left.
  • Zombie compute measures milliseconds of downstream work spent on responses the client already abandoned.
  • The blackhole drill contrasts a configured connectTimeout with the OS default ~75-second SYN retransmission.
Interview Round Script

Always separate connection timeout from read timeout — interviewers listen for that distinction. Then raise the ceiling: every hop needs a total budget smaller than its caller's, and the real pro move is deadline propagation via gRPC context or RequestDeadline headers so downstreams stop wasted work. Quote production numbers: 200-500 ms connect, 1-3 s read, tuned off p99.9.

Key Trade-Offs

Aggressive timeouts protect threads and latency SLOs but abort legitimately slow requests if set below real p99.

Related Curriculum Chapter

Timeout Design Patterns: Connection vs Read Timeouts

Read Full Chapter Blueprint

Explore More Interactive Labs

View All 280 Labs