TOPIC #190Intermediate 10 min read

Back-of-the-Envelope Estimation Techniques

CSD
CompleteSystemDesign Editorial
Report an issue
Key takeawayCore Architecture Summary

Master rapid capacity sizing and mental arithmetic: Powers of two vs powers of ten, the 100,000 seconds/day rule, Jeff Dean's latency numbers, and converting business scale into CPU, RAM, disk, and bandwidth requirements.

Key Glossary Concepts in this TopicAll Glossary Terms
Interactive Lab · 📐 100k Seconds/Day EstimatorFull lab guide

100,000-Seconds-a-Day Mental Math Trainer

Convert business scale into QPS, bandwidth, and storage the way interviewers expect.

Jeff Dean latency magnifier

The slower op is 750,000×× slower.

Whiteboard output

Average QPS3,000
300M ÷ 100,000 s
Peak QPS (×2)6,000
size for peak, not average
Throughput6.0 MB/s
= 48 Mbps (×8 bits)
Daily new data600.0 GB
×3 replicas +30% index → 2.34 TB/day physical
💡 Instant benchmarks: 1M/day ≈ 10 QPS · 100M/day ≈ 1,000 QPS · 1B/day ≈ 10,000 QPS. Bytes for storage, bits for NICs: 25 MB/s = 200 Mbps.
Powers of 2 vs 10: 2^10 ≈ 1 KB · 2^20 ≈ 1 MB · 2^30 ≈ 1 GB · 2^40 ≈ 1 TB. Round aggressively — order of magnitude is the deliverable, not decimals.

Back-of-the-Envelope Estimation Pipeline & Latency Hierarchy 📐

Systematic conversion from business DAU metrics to hardware capacity planning constraints alongside critical latency baselines.

Back-of-the-Envelope Estimation Pipeline & Latency Hierarchy 📐
100%
Touchpad: Pinch to zoom • Drag to pan
Rendering visual architecture flowchart...

01.The Philosophy of System Design Estimations

Back-of-the-envelope estimations are not intended to yield exact accounting precision down to decimal places. Rather, their purpose is to determine the order of magnitude (O(10^x)) of a system's physical constraints:

  • Are we designing for 100 QPS (single node) or 100,000 QPS (distributed cluster)?
  • Is 5-year storage 50 GB (single PostgreSQL database) or 50 PB (sharded S3 data lake)?
  • Is egress bandwidth 10 Mbps (commodity network card) or 100 Gbps (dedicated CDN edge routing)?

By rapidly bounding the problem space, you immediately discard non-viable architectural designs (e.g., attempting in-memory joins on 50 TB of data) and justify distributed components (caching, sharding, message queues) before drawing boxes.

02.Core Mental Math Constants & Approximations

The 100,000 Seconds per Day Rule

A day has 24 × 60 × 60 = 86,400 seconds. In mental calculations, round this up to 100,000 (10^5) seconds:

Average QPS = \frac{Total Daily Requests}{100,000}

Peak QPS = Average QPS × 2 (Standard Enterprise) \quad or \quad Average QPS × 3 to 5 (Flash Sales / Consumer)

Instant QPS Benchmarks:

  • 1 Million requests/day = \frac{10^6}{10^5} = 10 QPS (Peak ≈ 20-30 QPS)
  • 100 Million requests/day = \frac{10^8}{10^5} = 1,000 QPS (Peak ≈ 2,000-3,000 QPS)
  • 1 Billion requests/day = \frac{10^9}{10^5} = 10,000 QPS (Peak ≈ 20,000-30,000 QPS)

Powers of 2 vs Powers of 10 Approximation

Power of 2Exact ValuePower of 10 ApproxMetric Prefix
2^{10}1,02410³ (Thousand)1 Kilobyte (KB)
2^{20}1,048,57610^6 (Million)1 Megabyte (MB)
2^{30}1,073,741,82410^9 (Billion)1 Gigabyte (GB)
2^{40}1,099,511,627,77610^{12} (Trillion)1 Terabyte (TB)
2^{50}1,125,899,906,842,62410^{15} (Quadrillion)1 Petabyte (PB)

03.Latency Hierarchy: Jeff Dean's Numbers Every Engineer Must Know

Understanding the physics of latency across the computer storage and network hierarchy dictates why caching, indexing, and locality exist:

OperationLatencyScaled to Human Perspective (1ns = 1 sec)
L1 CPU Cache Reference0.5 - 1 ns1 second
Branch Mispredict3 ns3 seconds
L2 CPU Cache Reference3 - 4 ns4 seconds
Mutex Lock / Unlock17 ns17 seconds
Main Memory Reference (DRAM)100 ns1.5 minutes
Compress 1KB with Snappy2,000 ns = 2 µs33 minutes
Read 1 MB sequentially from Memory3 µs50 minutes
Read 4 KB random from NVMe SSD20 - 50 µs8 hours
Read 1 MB sequentially from SSD200 µs2.3 days
Datacenter LAN Roundtrip500 µs = 0.5 ms5.8 days
Send 1 MB over 10 Gbps network800 µs = 0.8 ms9.3 days
Read 1 MB sequentially from HDD (Disk)2,000 µs = 2 ms23 days
Disk Seek (Spinning Spindle)10 ms4 months
Cross-Continent Roundtrip (NYC to London)70 - 80 ms2.5 years
Global Trans-Pacific (NYC to Tokyo)160 - 200 ms6 years

Architectural Insights from Latency Physics:

  • Memory vs SSD: DRAM is 200-500× faster than NVMe SSDs for random reads.
  • Disk Seeks: Random disk seeks on spinning magnetic platters are 100,000× slower than RAM, explaining why LSM-Trees use sequential log writes.
  • Speed of Light: Network roundtrips across continents (100+ ms) dwarf internal server processing times (< 5 ms), mandating regional edge PoPs and CDNs.

04.Critical Safety Margins & Sizing Overheads

When estimating hardware infrastructure from raw data math, real-world systems incur structural overheads that must be factored in:

  1. Replication Factor Multiplier (3×): Enterprise distributed stores (HDFS, Kafka, Cassandra, Ceph) maintain 3 replicas across different availability zones or failure domains. Always multiply persistent raw storage by 3×.
  2. Database Index & Metadata Overhead (+30\% - 50\%): B-Tree indexes, primary key lookups, and table metadata (PostgreSQL MVCC vacuum headers) add 30\% - 50\% storage on top of raw payload bytes.
  3. Network Bit/Byte Conversion Trap: Network capacity is measured in Bits per second (bps, Gbps), while storage/memory is measured in Bytes per second (B/s, GB/s). Always multiply Bytes by 8 to obtain network line rates:

Bandwidth (Gbps) = Throughput (GB/s) × 8

  1. Headroom Buffer: Production systems should run at ≤ 65-70\% CPU and disk capacity to absorb sudden traffic surges, garbage collection pauses, and failover re-balancing.

Architectural Trade-offs & Production Realities

Architectural Advantages

  • Allows architects and candidates to ground complex architectural designs in hard physics within 3 minutes
  • Quickly eliminates unscalable approaches (e.g. attempting to scan 100TB tables synchronously on HTTP requests)

Trade-offs & Constraints

  • Over-focusing on exact arithmetic during interviews wastes precious time needed for high-level distributed systems trade-offs
  • Assuming average load instead of peak load leads to severe under-provisioning during live flash events
Production Implementation in Big Tech
Google / Meta• Production Cluster Allocation Sizing

Google and Meta engineering teams require formal Back-of-the-Envelope Capacity Estimations in every Design Document (Design Doc) before provisioning compute allocations, datacenter power quotas, and cross-region network backbone bandwidth for new global services.

Staff+ Engineering Takeaways

  • Use 100,000 seconds per day for instant, clean mental division ($1\text{M}/\text{day} = 10\text{ QPS}$).
  • RAM is 500x faster than SSD and 50,000x faster than network roundtrips.
  • Multiply persistent storage by 3x for cross-AZ replication and add 30-50% for indexing overhead.
  • Always distinguish between Bytes ($B$) for storage and Bits ($b$) for network line bandwidth ($1\text{ MB/s} = 8\text{ Mbps}$).

Topic Knowledge Check

Exercise 1 of 3 • Test your architectural comprehension.

Exercise 1 of 30 answered
1

A photo-sharing service processes 300 million API read requests per day. Using standard back-of-the-envelope approximations, what is the estimated Average QPS and recommended Peak QPS capacity?

Rate This Architecture ChapterFeedback & Rating

How clear and actionable was this distributed systems breakdown?

Interactive Engineering Workbenches: