Home/Labs/Discord Scylla Tail Latency
All 280 Labs
INTERACTIVE LAB💬

Discord ScyllaDB Tail-Latency Lab (Interactive)

Swap JVM Cassandra for C++ ScyllaDB and bucket hot channels to flatten the p99. Store trillions of messages on a thread-per-core storage ring: flip engines to feel garbage-collection spikes, size (channel_id, time_bucket) partitions, and count the nodes each engine needs.

Discord Storage Ring & Tail-Latency Lab

Run billions of daily messages through JVM Cassandra vs C++ ScyllaDB, and see how (channel_id, time_bucket) partitioning stops runaway hot rows.

p50 message read1.8 ms
p99 message read4 ms
Garbage collectionNone — native C++ memory control
Storage nodes required2 ScyllaDB nodes
Hottest partition rows200,000 over 10 days
Partition size vs 100 MB limit50.0 MB — bounded
Thread-per-core routing with zero GC keeps p99 under 5 ms and time buckets cap partition growth at 50 MB — trillions of messages, no spinners.

How It Works Under the Hood

Discord's message store went MongoDB to Cassandra to ScyllaDB as volume raced past trillions of rows. JVM Cassandra was masterless and linearly scalable, but garbage-collection stop-the-world pauses spiked p99 reads past a second — loading spinners in a real-time chat. ScyllaDB's C++ rewrite binds one thread per CPU core with no GC, dropping tails under five milliseconds on a fraction of the fleet. Beside the storage ring, Elixir BEAM processes hold guild state in RAM with Rust NIFs crunching permission bitmasks at native speed.

Core Architectural Principles

  • Tail-latency physics: JVM GC pauses inject 1,000-5,000 ms spikes into p99 reads; thread-per-core Seastar removes the pause source entirely.
  • Time-bucketed partition keys (channel_id, 10-day bucket) bound partition size under the ~100 MB LSM danger threshold.
  • Hot-path split: Elixir actors for concurrency and guild state, Rust NIFs for CPU-bound permission and sorting work at native speed.
Interview Round Script

When asked why Discord left Cassandra, answer in tail-latency terms: "Average was fine; JVM GC spikes wrecked p99, and chat is a p99 product." Then design the partition key before being prompted: composite (channel_id, bucket) with 10-day windows. Finally, map layers correctly — Elixir owns real-time guild state, ScyllaDB owns durability — interviewers probe that separation.

Key Trade-Offs

ScyllaDB buys deterministic single-digit-millisecond tails with a smaller C++ footprint but demands disciplined partition-key design to avoid wide-row scans.

Related Curriculum Chapter

Discord: Scaling Elixir, Rust, & ScyllaDB for Millions of Concurrent Users

Read Full Chapter Blueprint

Explore More Interactive Labs

View All 280 Labs