100x Scaling Question Lab (Interactive)
Viral growth hits: shard the database, pool connections, buffer writes, add regions until nothing breaks first. Scale traffic from 1x to 100x and watch which component saturates first, then evolve the architecture with consistent-hash sharding, PgBouncer, Kafka write buffering, and multi-region active-active deployment.
100x Scaling Stress Test
Viral growth hits: find what breaks first, then evolve with sharding, pooling, buffering, and regions - without over-engineering the 1x baseline.
Architecture absorbs 1x - 4-pillar pitch ready
Deliver the executive pitch: 1) horizontal sharding on user_id, 2) L1+L2 multi-tier caching, 3) Kafka-buffered non-critical writes, 4) 1 region active-active with async CDC replication and conflict resolution (LWW/CRDT).
How It Works Under the Hood
The classic finale separates candidates who know what breaks from those who say "everything crashes." Stateless app nodes scale effortlessly behind load balancers, but a single write primary saturates at roughly 25k transactions per second because WAL disk serialization and row locks cannot be cloned - they must be partitioned. Connection counts explode with the app fleet, the 80/20 hot set outgrows one Redis box, and at 50x, speed-of-light latency makes one region indefensible. This lab models each ceiling numerically and only declares the architecture ready when shards, pools, buffers, and regions absorb the multiplier.
Core Architectural Principles
- Per-shard write ceiling (~25k TPS), pooled connection budgets, and 80/20 cache growth recomputed at every multiplier.
- Shard key choice on user_id keeps user-scoped transactions local; cross-shard reads move to CDC-fed search clusters.
- Multi-region active-active collapses global p99 from 170 ms to sub-30 ms with asynchronous cross-region replication.
Answer the 100x question with a pinpoint: "the single primary database write path fails first." Then deliver the four-pillar pitch - horizontal sharding with consistent hashing, L1+L2 caching, Kafka-buffered non-critical writes, and multi-region active-active with GeoDNS. Note the operational costs of sharding and cross-region conflict resolution to show you are not overselling, and never pre-build 100x architecture for a 10k-user prompt.
Evolutionary scaling proves the baseline extends to hyperscale, but sharding and active-active regions add rebalancing, cross-shard query, and conflict-resolution complexity.