Home/Labs/Live Shard Migration Runway
All 280 Labs
INTERACTIVE LAB🔄

Shard Rebalancing Migration Lab (Interactive)

Copy 10 TB to new shards under live writes and watch backfill, CDC backlog, validation, and cutover unfold. Step through the four-stage zero-downtime migration: snapshot backfill with throttling, Debezium CDC catch-up, shadow validation, atomic cutover.

Zero-Downtime Shard Rebalancing Pipeline

Move 10 TB to new shards while live traffic writes: backfill, CDC catch-up, shadow validation, cutover.

IdleStage 1: Historical BackfillStage 2: CDC Catch-UpStage 3: Shadow ValidationStage 4: Atomic CutoverComplete

Data copied

0 / 10,000 GB

CDC queue

0 ops

lag ≈ 0 ms behind live

Shadow validation

0%

peak residual backlog: 0 ops

Downtime so far

0 ms

cutover = in-memory routing flip

The invariant this models: the backfill can never finish “perfectly” because live writes keep mutating the source. Only a continuously streaming CDC channel (Debezium + Kafka replaying WAL at 25k ops/s) can drain the queue to zero lag, and only 100% shadow checksum equality unlocks the cutover.

How It Works Under the Hood

Rebalancing terabytes across a cluster serving fifty thousand requests per second cannot pause the business. Stage one copies a consistent snapshot with I/O throttling to protect production disk bandwidth. Stage two streams every concurrent mutation from the source WAL through Kafka to the destination until replication lag drains to milliseconds. Stage three shadow-reads both clusters comparing checksums, and stage four flips an in-memory routing table, keeping old shards as rollback standbys.

Core Architectural Principles

  • Live writes accumulate as a CDC backlog the destination must drain faster than they arrive.
  • Migration throttles automatically when write pressure or backlog exceeds safe thresholds.
  • Cutover is atomic: routing pointers flip in etcd/ZooKeeper, so customer-visible downtime stays zero.
Interview Round Script

Walk the four stages in order and name the tools: snapshot backfill, Debezium plus Kafka catch-up, shadow checksum reads, atomic proxy cutover. Emphasize the drain-rate inequality, CDC must outpace live writes or lag never closes, and mention Vitess VReplication automating this for MySQL fleets. Note the 2x temporary storage cost.

Key Trade-Offs

Perfect uptime and verified integrity versus duplicated storage, CDC pipelines, and heavy orchestration machinery.

Related Curriculum Chapter

Rebalancing Shards & Data Migration

Read Full Chapter Blueprint

Explore More Interactive Labs

View All 280 Labs