Load Balancer Failover Lab (Interactive)
Kill a backend or the Active LB and measure VRRP failover and health-check detach times. Drive an active-passive HA pair with a shared Virtual IP. Tune probe intervals, failure thresholds, and heartbeats, then compare L4 TCP versus L7 HTTP health checks.
Eliminating the Load Balancer SPOF
Crash a backend or the Active LB and watch VRRP failover, health-check detection, and connection draining react.
Standby promotes after 3 missed VRRP heartbeats + Gratuitous ARP.
GET /healthz expecting 200 OK verifies true readiness.
10.0.1.11:8080
CPU: 32%
10.0.1.12:8080
CPU: 28%
10.0.1.13:8080
CPU: 41%
10.0.1.14:8080
CPU: 37%
SCENARIO TRACE
› 1,000 req/s balanced across app1, app2, app3, app4. VIP stable on Primary LB.
Tuning knobs
Why L7 checks exist
A TCP handshake only proves the socket is open. Switch the probe to L4 TCP and press Hang App Server 3: a deadlocked process keeps answering SYN-ACK, so the LB never detaches it and a slice of live traffic errors forever.
Tight VRRP heartbeats (<200ms) give sub-second VIP failover; balance probe interval against check noise on the fleet.
How It Works Under the Hood
A load balancer in front of fifty servers is itself a Single Point of Failure: one crashed appliance takes the whole system down. Production systems neutralize it with an Active-Passive VRRP/Keepalived pair sharing a floating VIP — the standby promotes after three missed heartbeats and claims the VIP with a Gratuitous ARP broadcast in under a second. Meanwhile health checks decide how fast a dying backend detaches: L4 TCP probes only prove the socket answers SYN, while L7 GET /healthz checks verify true application readiness, and connection draining lets in-flight requests finish before a server retires during a deploy.
Core Architectural Principles
- VRRP active-passive pair: standby claims the shared VIP via Gratuitous ARP after 3 missed heartbeats.
- L4 TCP handshake checks miss deadlocked processes; L7 /healthz checks measure real readiness (interval × fall threshold).
- Connection draining stops new requests while in-flight connections finish for zero-downtime deploys.
When asked "isn't your load balancer a single point of failure?", answer with two layers: Active-Passive VRRP/Keepalived with a floating VIP for bare metal (sub-second GARP failover), and BGP Anycast with ECMP for hyperscale clouds. Then mention connection draining as the mechanism behind zero-downtime rolling restarts.
Tight heartbeat and probe intervals detect failures faster but add control-plane traffic noise and false-positive detachments.