Multi-Region Active-Active vs Active-Passive Lab (Interactive)
Kill the primary region and count downtime minutes, lost transactions, split-brain divergence, and standby cost for each topology. Model cold, warm, and hot standby versus Active-Active against RTO/RPO targets, replication lag, fencing tokens, and write latency.
RTO / RPO Regional Failover Drill
Kill a region and count the downtime minutes and lost transactions the topology permits.
How It Works Under the Hood
Disaster recovery is a two-number contract: RTO bounds downtime, RPO bounds data loss in time. Async replication gives a hot standby RTO near a minute and an RPO equal to its lag — at 5,000 tx/sec, two lost seconds is 10,000 transactions. Active-Active approaches zero on both only with multi-master conflict resolution: CRDTs that converge regardless of arrival order, or Spanner TrueTime commit-wait over atomic-clock-bounded intervals, and it pays for that with cross-region synchronous write latency and 50-100% standing-cost premium. Fencing tokens prevent the both-regions-claim-primary split-brain corruption.
Core Architectural Principles
- RPO × write QPS = transactions lost per outage; async lag slider is literally your data-loss budget.
- Fencing epochs issued by etcd/ZooKeeper make the demoted primary reject writes after promotion.
- Active-Active evacuation works when each region idles at 60-70% and can absorb a neighbor's surge.
Before choosing anything, ask "what are the RTO and RPO targets?" Then justify the topology from the answer: minutes-and-seconds for Active-Passive e-commerce, near-zero only where CRDTs or distributed SQL exist. Drop fencing tokens and split-brain to survive the failover follow-up, and Chaos Kong to show evacuation is a drill, not a hope.
Active-Passive keeps simple ACID single-writer semantics with minute-scale recovery, while Active-Active reaches zero RTO/RPO only by accepting multi-master conflict complexity and paying double for idle headroom.