Home/Labs/DNS Global Failover
All 280 Labs
INTERACTIVE LAB🛰️

DNS Global Load Balancing Lab (Interactive)

Crash a region and watch TTL-cached resolvers keep users pinned to the outage while your record already flipped. Choose Round Robin, Geo, Latency, or Failover routing, set the record TTL, and measure misrouted traffic drain, evacuation time, and authoritative DNS query load.

Global DNS Routing & TTL Failover

Kill the Tokyo region and watch resolver-cached records keep sending users into the outage until TTL expires.

us-east-1 (Virginia)

A record → 198.51.100.10

500K users · +25ms

HEALTHY

eu-central-1 (Frankfurt)

A record → 203.0.113.50

300K users · +30ms

HEALTHY

ap-northeast-1 (Tokyo)

A record → 192.0.2.99

200K users · +35ms

HEALTHY

Traffic still stuck on dead Tokyo — resolver cache drain (TTL 300s)

+0 min
100%
+2 min
100%
+5 min
100%
+10 min
100%
+20 min
100%
+40 min
100%
+80 min
100%

Authoritative flip happens at +15s (detection), but propagation = TTL. A Tokyo user querying through 8.8.8.8 holds the stale IP until its cached entry expires; some ISP resolvers “clamp” TTLs to hours.

DNS knobs

Current policy: latency — lowest measured RTT per user ISP.

Users still misrouted0K (100%)
Time to 90% evacuate285s
Auth DNS QPS6,667
Misrouted req share0.0%

The TTL tug-of-war

Low TTL (30s): fast failover, ×10 DNS query volume.

High TTL (86400s): cheap lookups, hours-long failover.

Answer: pair GSLB with BGP Anycast for sub-second, then DNS handles region choice.

How It Works Under the Hood

Global Server Load Balancing steers traffic at the very first step: an intelligent authoritative nameserver like Route 53 inspects the resolver's EDNS Client Subnet, live latency measurements, and health-check state before returning an IP. Its Achilles heel is distributed caching: resolvers honor the record TTL, so after a Tokyo outage the authoritative answer flips within ~15 seconds of detection, but every user whose ISP cached the old record keeps hammering the dead region until expiry — some resolvers even clamp TTLs to hours. The fix is operational: TTLs of 60 seconds or less for critical endpoints, multi-value answers, and pairing DNS with BGP Anycast for sub-second failover.

Core Architectural Principles

  • Five policies: weighted round-robin, Geo/EDNS Client Subnet, latency-based, active-passive failover, multi-value answers.
  • Health checkers flip the record ~15s after detection, but traffic drain follows the TTL curve, not the flip.
  • Lower TTL = faster evacuation but higher authoritative query QPS; TTL clamping by rogue resolvers defeats both.
Interview Round Script

When designing multi-region, give the entry pipeline: Route 53 GSLB to regional anycast L4, then L7 gateways. Show the TTL tradeoff quantitatively — time-to-evacuate ≈ detection + TTL — and close with the hybrid answer: low-TTL DNS plus BGP Anycast for sub-second failover, since DNS alone cannot be instantaneous.

Key Trade-Offs

DNS steering spans regions and clouds with no proxy bottleneck, but TTL caching makes failover gradual and it sees no per-server CPU or memory load.

Related Curriculum Chapter

DNS-Based Load Balancing & Global Traffic Management

Read Full Chapter Blueprint

Explore More Interactive Labs

View All 280 Labs