Magic Pocket Erasure Coding Lab (Interactive)
Migrate exabytes off S3: Reed-Solomon shards versus 3x replication, chunk by chunk. Price Dropbox's custom storage: dedup content-addressed 4 MB chunks, encode them as 9+3 or 12+4 Reed-Solomon shards, kill racks, and measure the bandwidth delta sync saves on one edited file.
Magic Pocket Durability & Delta Sync Ledger
Trade 3x replication against Reed-Solomon erasure coding on exabytes of 4 MB chunks, then measure what delta sync saves a single edited file.
How It Works Under the Hood
Serving five hundred petabytes from S3 was economically absurd for Dropbox, so Magic Pocket moved the exabyte scale onto custom hardware with erasure coding instead of replication. Each deduplicated chunk is split into k data shards plus m parity shards — 9+3 tolerates any three rack failures at 33% overhead versus 200% for 3x replication. Durability comes from reconstruct-on-loss across racks, and sync speed comes from the delta protocol: the client's local SQLite hash index tells the server which 4 MB blocks actually changed, so a edited one-gigabyte file re-uploads kilobytes, not gigabytes.
Core Architectural Principles
- Erasure coding overhead: k+m shards store data at m/k extra bytes while surviving any m simultaneous shard losses.
- Content-addressed dedup: SHA-256 chunk hashes mean identical blocks across users are stored exactly once, before coding.
- Delta sync: a client-side hash manifest reduces a full-file re-upload to only the 4 MB chunks whose digests changed.
For object-storage design, quantify the coding decision: "3x replication costs 300% for two-failure tolerance; RS 9+3 costs 133% for three." Mention reconstruction bandwidth as the hidden cost — losing a shard rereads k shards to rebuild it. Then cover sync semantics: block-level dedup plus delta hashing is what makes Dropbox feel instant on huge files.
Erasure coding slashes storage cost but pays reconstruction bandwidth and latency; custom hardware beat S3 at scale only because Dropbox volume justified building it.