Twitter Hybrid Fan-Out Lab (Interactive)
Flip celebrity accounts from push to pull and watch writes per second collapse at scale. Fan out tweets to follower timelines eagerly or hydrate them at read time, tune the celebrity threshold, and size Redis sorted-set timelines against delivery latency and memory.
Twitter Fan-Out on Write vs Read
Push every tweet to follower Redis timelines or pull celebrities at read time — the hybrid 50k threshold decides between the two and sizes the DRAM bill.
How It Works Under the Hood
Twitter timeline generation is the canonical fan-out-on-write case. When a normal account tweets, the pipeline pushes the tweet ID into each follower's Redis sorted set of the most recent 800 IDs, making Home Timeline reads a merge of precomputed sets plus an asynchronous MySQL backfill. Celebrities with millions of followers break that model, so Twitter went hybrid: eager push for regular accounts, pull-on-read for the huge-follower cohort, all fronted by Snowflake IDs that sort globally by time despite being generated on thousands of machines.
Core Architectural Principles
- Push cost per tweet scales with follower count: the fan-out write amplification is exactly the celebrity problem quantified.
- Hybrid threshold: accounts above roughly 50,000 followers are fetched at read time, capping worst-case writes per tweet.
- Timeline memory: 800 tweet IDs per follower in Redis sorted sets fixes the DRAM envelope per user of the read cache.
In timeline rounds, never pick one strategy. Say: "Fan-out on write for regular users, hybrid pull for celebrities, threshold around 50,000 followers." Then bring Snowflake: 41-bit timestamp, 10 machine bits, 12 sequence — monotonic without a central counter. Quoting the 800-ID Redis cap shows you know timelines are an LRU cache, not a database copy.
Eager fan-out buys instant timeline reads at write amplification; pull-on-read saves celebrity writes but pays merge cost on every home-screen load.