Estimating QPS, Storage, & Bandwidth: End-to-End Walkthrough
Work through a comprehensive, rigorous end-to-end capacity planning model: Sizing a 500M DAU global social platform from QPS and 5-year multi-tier storage to network egress bandwidth and Redis cluster RAM sizing.
500M-DAU End-to-End Capacity Pipeline
QPS → 5-year storage → line rates → 80/20 RAM cache → concrete node, shard, and pod counts.
Step 1 — Traffic
Steps 2-3 — Storage & bandwidth
Step 4 — Fleet sizing
Complete 500M DAU End-to-End Capacity Estimation Pipeline 📊
A structured, rigorous derivation from user activity metrics to compute instances, persistent storage tiers, network line rates, and in-memory cache allocations.
01.Problem Statement & Baseline Assumptions
Let us design the capacity and infrastructure requirements for a global social platform (e.g., a Twitter/X or Instagram hybrid) with the following product scale specifications:
- Daily Active Users (DAU):
500 Million - User Activity (Writes):
100 Million new posts/day(0.2 posts/user/day).100\%contain text & metadata (500 bytes).10\%of posts include an image (2 MBaverage size).
- User Activity (Reads):
10 Billion feed requests/day(Each active user opens the feed20 times/day).- Each feed request returns a page of 20 post summaries
≈ 10 KBresponse payload.
- Each feed request returns a page of 20 post summaries
- Read-to-Write Ratio:
10 Billion reads : 100 Million writes = 100:1 (Heavily Read-Heavy).
02.Mathematical Step-by-Step Capacity Derivation
Step 1: Traffic Estimation (QPS Modeling)
Using the 100,000 seconds/day mental approximation:
Average Write QPS = \frac{100,000,000 writes}{100,000 sec} = 1,000 Write QPS
Peak Write QPS = 1,000 × 2 = 2,000 Peak Write QPS
Average Read QPS = \frac{10,000,000,000 reads}{100,000 sec} = 100,000 Read QPS
Peak Read QPS = 100,000 × 2 = 200,000 Peak Read QPS
Step 2: Storage Sizing (Daily & 5-Year Projections)
- Metadata & Text Database Storage:
- Daily volume:
100M posts × 500 bytes = 50 GB/day. - 5-Year persistent volume:
50 GB/day × 365 days × 5 years = 91,250 GB ≈ 91.25 TB. - With
3×cross-AZ replication+ 30\%index overhead:
- Daily volume:
91.25 TB × 3 × 1.30 = 355.8 TB Raw Disk Capacity
- Blob / Object Storage (Photos & Media):
10\%of100M posts = 10M images/day.- Daily media ingestion:
10M images × 2 MB = 20,000,000 MB = 20 TB/day. - 5-Year media volume:
20 TB/day × 365 × 5 ≈ 36.5 Petabytes (PB). - Architecture decision: Store images in Amazon S3 / Google Cloud Storage with automated lifecycle policies transitioning objects older than 90 days to S3 Glacier Deep Archive (
95\%cost savings).
Step 3: Network Bandwidth Estimation
-
Ingress Bandwidth (Writes):
- Media Ingress:
10M images/day × 2 MB = 20 TB/day = \frac{20 × 10^{12} bytes}{10^5 sec} = 200 MB/s. - Text Ingress:
1,000 QPS × 500 B = 0.5 MB/s. - Total Ingress Line Rate:
200.5 MB/s × 8 bits/byte = 1.6 Gbps.
- Media Ingress:
-
Egress Bandwidth (Reads):
- API Feed Egress:
100,000 Read QPS × 10 KB = 1,000,000 KB/s = 1 GB/s. - In Line Rate:
1 GB/s × 8 bits/byte = 8 Gbps Egress. - With a Global CDN (Cloudflare/CloudFront) caching
90\%of feed assets and media at the edge, origin datacenter egress drops from8 Gbpsto800 Mbps, protecting backend API infrastructure.
- API Feed Egress:
Step 4: Memory Cache Sizing (80/20 Pareto Working Set)
- Daily API read volume:
10 Billion views × 10 KB = 100 TB read data/day. - Applying the 80/20 Pareto Rule,
20\%of the unique daily content generates80\%of read traffic. - Required in-memory cache size:
RAM Cache Size = 100 TB × 0.20 = 20 TB RAM
03.Hardware & Cluster Sizing: Translating Math to Server Fleets
A senior system design interview requires connecting mathematical storage and bandwidth numbers to concrete physical node fleet counts:
-
API Application Fleet Sizing:
- A single optimized stateless Go/Java/Node.js API container handles
~ 2,000 QPSunder normal CPU load. - Peak Read Traffic =
200,000 QPS. - Required API instances =
\frac{200,000 Peak QPS}{2,000 QPS/instance} = 100 API Pods(Provision 130 pods for30\%headroom and redundancy).
- A single optimized stateless Go/Java/Node.js API container handles
-
Redis In-Memory Caching Cluster:
- Total RAM required =
20 TB. - Standard AWS memory-optimized node (
r6g.4xlarge):128 GB RAM. - Usable RAM per node (leaving
25\%for Redis copy-on-write BGSAVE overhead)≈ 96 GB. - Cluster Node Count =
\frac{20,000 GB}{96 GB/node} ≈ 208 Redis Primary Nodes(+208replicas across AZs for high availability).
- Total RAM required =
-
Database Sharding Fleet:
- 5-Year Metadata Storage =
91.25 TB. - To keep PostgreSQL/MySQL B-Tree index scans fast and backup restore times under 30 minutes, limit each database shard size to
≤ 2 TB SSD. - Shard Count =
\frac{91.25 TB}{2 TB/shard} ≈ 46 Shards(Round to64 shardsfor clean power-of-2 consistent hashing partition rings).
- 5-Year Metadata Storage =
04.The Capacity Estimation Summary Card
| Metric Dimension | Average Load | Peak Load (2x) | 5-Year Cumulative | Infrastructure Footprint |
|---|---|---|---|---|
| Write QPS | 1,000 ops/sec | 2,000 ops/sec | 182.5 Billion records | Kafka Ingestion Buffer (6 brokers) |
| Read QPS | 100,000 ops/sec | 200,000 ops/sec | — | 130 Stateless API Pods |
| Relational Storage | 50 GB/day | — | 91.25 TB (355 TB with 3x repl + index) | 64 Database Shards (2TB each) |
| Media Blob Storage | 20 TB/day | — | 36.5 Petabytes | S3 Object Store + Glacier Lifecycle |
| Network Egress | 1 GB/s (8 Gbps) | 2 GB/s (16 Gbps) | — | Edge CDN with 90% cache offload |
| Memory Cache (RAM) | — | — | — | 208 Redis Nodes (20 TB total working set) |
Architectural Trade-offs & Production Realities
Architectural Advantages
- Gives concrete, defensible numbers for every architectural component (shards, pods, cache nodes, object tiers)
- Reveals critical scaling bottlenecks early (e.g. realizing media storage requires Petabyte-scale object storage rather than relational DBs)
Trade-offs & Constraints
- Assumptions may shift drastically during actual product growth (e.g. video introduction vs pure text)
- Failure to account for network line rate conversion ($8\times$) leads to severe network interface card (NIC) saturation
Twitter engineers use strict 5-year storage projections and 80/20 cache sizing formulas to determine SSD partition allocations across Manhattan distributed key-value stores and to size memory footprints across tens of thousands of Twemcache (Memcached) instances.
Staff+ Engineering Takeaways
- Always follow the standard 5-step capacity sizing sequence: QPS -> Storage -> Bandwidth -> Memory -> Node Counts.
- Use the 80/20 rule to size the RAM caching tier (cache 20% of daily read volume).
- Divide cumulative 5-year storage by 2TB to calculate the number of database shards required.
- Apply CDN edge caching to offload 85-95% of egress bandwidth away from origin microservices.
Topic Knowledge Check
Exercise 1 of 3 • Test your architectural comprehension.
A photo platform receives 50 million photo uploads per day, with an average photo size of 1.5 Megabytes. Approximately how much raw object storage will be consumed over 5 years (ignoring compression)?
How clear and actionable was this distributed systems breakdown?