DNS Global Load Balancing Lab (Interactive)
Crash a region and watch TTL-cached resolvers keep users pinned to the outage while your record already flipped. Choose Round Robin, Geo, Latency, or Failover routing, set the record TTL, and measure misrouted traffic drain, evacuation time, and authoritative DNS query load.
Global DNS Routing & TTL Failover
Kill the Tokyo region and watch resolver-cached records keep sending users into the outage until TTL expires.
us-east-1 (Virginia)
A record → 198.51.100.10
500K users · +25ms
HEALTHY
eu-central-1 (Frankfurt)
A record → 203.0.113.50
300K users · +30ms
HEALTHY
ap-northeast-1 (Tokyo)
A record → 192.0.2.99
200K users · +35ms
HEALTHY
Traffic still stuck on dead Tokyo — resolver cache drain (TTL 300s)
Authoritative flip happens at +15s (detection), but propagation = TTL. A Tokyo user querying through 8.8.8.8 holds the stale IP until its cached entry expires; some ISP resolvers “clamp” TTLs to hours.
DNS knobs
Current policy: latency — lowest measured RTT per user ISP.
The TTL tug-of-war
Low TTL (30s): fast failover, ×10 DNS query volume.
High TTL (86400s): cheap lookups, hours-long failover.
Answer: pair GSLB with BGP Anycast for sub-second, then DNS handles region choice.
How It Works Under the Hood
Global Server Load Balancing steers traffic at the very first step: an intelligent authoritative nameserver like Route 53 inspects the resolver's EDNS Client Subnet, live latency measurements, and health-check state before returning an IP. Its Achilles heel is distributed caching: resolvers honor the record TTL, so after a Tokyo outage the authoritative answer flips within ~15 seconds of detection, but every user whose ISP cached the old record keeps hammering the dead region until expiry — some resolvers even clamp TTLs to hours. The fix is operational: TTLs of 60 seconds or less for critical endpoints, multi-value answers, and pairing DNS with BGP Anycast for sub-second failover.
Core Architectural Principles
- Five policies: weighted round-robin, Geo/EDNS Client Subnet, latency-based, active-passive failover, multi-value answers.
- Health checkers flip the record ~15s after detection, but traffic drain follows the TTL curve, not the flip.
- Lower TTL = faster evacuation but higher authoritative query QPS; TTL clamping by rogue resolvers defeats both.
When designing multi-region, give the entry pipeline: Route 53 GSLB to regional anycast L4, then L7 gateways. Show the TTL tradeoff quantitatively — time-to-evacuate ≈ detection + TTL — and close with the hybrid answer: low-TTL DNS plus BGP Anycast for sub-second failover, since DNS alone cannot be instantaneous.
DNS steering spans regions and clouds with no proxy bottleneck, but TTL caching makes failover gradual and it sees no per-server CPU or memory load.