Shard Rebalancing Migration Lab (Interactive)
Copy 10 TB to new shards under live writes and watch backfill, CDC backlog, validation, and cutover unfold. Step through the four-stage zero-downtime migration: snapshot backfill with throttling, Debezium CDC catch-up, shadow validation, atomic cutover.
Zero-Downtime Shard Rebalancing Pipeline
Move 10 TB to new shards while live traffic writes: backfill, CDC catch-up, shadow validation, cutover.
Data copied
0 / 10,000 GB
CDC queue
0 ops
lag ≈ 0 ms behind live
Shadow validation
0%
peak residual backlog: 0 ops
Downtime so far
0 ms
cutover = in-memory routing flip
The invariant this models: the backfill can never finish “perfectly” because live writes keep mutating the source. Only a continuously streaming CDC channel (Debezium + Kafka replaying WAL at 25k ops/s) can drain the queue to zero lag, and only 100% shadow checksum equality unlocks the cutover.
How It Works Under the Hood
Rebalancing terabytes across a cluster serving fifty thousand requests per second cannot pause the business. Stage one copies a consistent snapshot with I/O throttling to protect production disk bandwidth. Stage two streams every concurrent mutation from the source WAL through Kafka to the destination until replication lag drains to milliseconds. Stage three shadow-reads both clusters comparing checksums, and stage four flips an in-memory routing table, keeping old shards as rollback standbys.
Core Architectural Principles
- Live writes accumulate as a CDC backlog the destination must drain faster than they arrive.
- Migration throttles automatically when write pressure or backlog exceeds safe thresholds.
- Cutover is atomic: routing pointers flip in etcd/ZooKeeper, so customer-visible downtime stays zero.
Walk the four stages in order and name the tools: snapshot backfill, Debezium plus Kafka catch-up, shadow checksum reads, atomic proxy cutover. Emphasize the drain-rate inequality, CDC must outpace live writes or lag never closes, and mention Vitess VReplication automating this for MySQL fleets. Note the 2x temporary storage cost.
Perfect uptime and verified integrity versus duplicated storage, CDC pipelines, and heavy orchestration machinery.