TOPIC #97Beginner 7 min read

Why Caching Matters: The Latency & Cost Multiplier

CSD
CompleteSystemDesign Editorial
Report an issue
Key takeawayCore Architecture Summary

Explore the physics of caching: RAM vs Disk latency differentials, database offloading, Pareto principle (80/20 rule), and egress cost reductions.

Key Glossary Concepts in this TopicAll Glossary Terms
Interactive Lab · ⚡ Cache Hit Ratio EconomicsFull lab guide

Cache Hit Ratio Economics Lab

See why a 99% hit ratio is 10× better than 90%: every point of misses is paid for in database cores.

Workload Inputs

Cache QPS (effective)
100.0k/s
baseline traffic
DB QPS with cache
5.0k/s
5.0% miss ratio
DB QPS no cache
100.0k/s
every read hits disk tier
Offload factor
20.0×
DB load removed by cache
Avg read latency
1.75 ms
hit 0.5ms / miss 25ms
DB nodes needed
1
vs 9 uncached
DB cluster cost
$1.2k/mo
saves $9.6k/mo
Hot-set RAM size
74.5 GB
20% of 200.00M keys
Latency & cost derivation
MISS_RATIO = 1 − 95.0% = 5.0%
DB_QPS = 100.0k × 5.0% = 5.0k/s
AVG_LATENCY = 95.0% × 0.5ms + 5.0% × 26ms = 1.75 ms
RAM is ~100ns, NVMe ~50µs, SQL query ~25ms — a 50,000× gap the cache bridges.
CACHE saves $9.6k/mo in DB nodes (200.00M keys × 2000B hot set = 74.5 GB RAM)

95% is the industry target zone. Dropping to 90% would push 10.0k QPS at the DB instead of 5.0k — often forcing a cluster upgrade.

Rule of thumb: one $1.2k/mo DB node handles ~12.0k read QPS. A small Redis node serves 100k+ QPS for a fraction of that — caching is a cost multiplier, not just a speed trick.

Database Offloading via Caching 🚀

95%+ of read traffic intercepted in RAM before hitting disk.

Database Offloading via Caching 🚀
100%
Touchpad: Pinch to zoom • Drag to pan
Rendering visual architecture flowchart...

01.The Physical Latency Differential: RAM vs Disk vs Network

At the physical hardware layer, the fundamental law of computer architecture dictates that speed is inversely proportional to capacity and distance:

  • L1/L2 CPU Cache: ~ 0.5 - 4 nanoseconds (32KB - 512KB).
  • Main Memory (DRAM): ~ 100 nanoseconds (8GB - 512GB).
  • NVMe Solid-State Storage (SSD): ~ 20 - 50 microseconds (20,000 - 50,000 ns, roughly 500× slower than RAM).
  • Relational SQL Database Query: ~ 5 - 50 milliseconds (5,000,000 - 50,000,000 ns, roughly 50,000× slower than RAM).
  • Cross-Datacenter Network Roundtrip: ~ 70 - 150 milliseconds.

By placing an in-memory cache (such as Redis or Memcached) in front of your database layer, you bridge this 50,000x physical speed gap, serving queries from DRAM at sub-millisecond latencies.

02.The Pareto Principle (80/20 Rule) & Cache Hit Ratio Economics

In nearly all consumer and enterprise web applications, read access patterns are heavily skewed according to the Pareto Principle (80/20 Rule):

  • 80\% of all read traffic targets 20\% of the data (e.g., top trending tweets, viral YouTube videos, active product catalog items, current session tokens).
  • By provisioning just enough RAM to hold this hot working set (20%), you can intercept 90\% to 99\% of all incoming queries.

The Cache Hit Ratio Formula:

Cache Hit Ratio = \frac{Cache Hits}{Cache Hits + Cache Misses} × 100\%

Why a 99% Hit Ratio is 10× Better than 90%:

  • At 100,000 QPS with a 90\% hit ratio, 10,000 QPS penetrates to the database.
  • At 100,000 QPS with a 99\% hit ratio, only 1,000 QPS reaches the database.
  • Improving hit ratio from 90\% to 99\% reduces database load by a factor of 10, preventing multi-thousand-dollar database cluster upgrades.

03.Architectural Value: Cost Reduction & Traffic Spikes

Beyond raw latency improvements, caching provides critical architectural protections:

  1. Database Scaling Cost Optimization: Scaling a PostgreSQL or MySQL primary database vertically (e.g., AWS db.r6i.32xlarge with 128 vCPUs and 1TB RAM) costs thousands of dollars per month. In contrast, a small Redis Cluster node delivers 100,000+ QPS for a fraction of the cost.
  2. Shock Absorber for Traffic Surges: During flash sales (Black Friday) or breaking news events, traffic can surge 10× within seconds. A well-designed cache absorbs this spike in memory without overwhelming database connection pools or triggering table lock cascades.
  3. Cloud Egress & Microservice Protection: Caching computed API responses at the edge (CDN) or API gateway eliminates downstream microservice RPC hops and reduces cross-AZ cloud network egress bills.

Architectural Trade-offs & Production Realities

Architectural Advantages

  • Delivers sub-millisecond read latency ($< 1\text{ms}$) by serving data directly from RAM
  • Massively offloads database CPU, I/O operations, and connection pool contention
  • Acts as a resilient shock absorber against sudden viral traffic spikes

Trade-offs & Constraints

  • Introduces data consistency challenges: cached data can become stale if underlying records mutate
  • Requires explicit invalidation logic, TTL policies, and memory eviction management
Production Implementation in Big Tech
Twitter / X• Timeline Caching

Twitter maintains billions of user home timelines in massive Redis clusters, serving millions of timeline views per second directly from memory.

Staff+ Engineering Takeaways

  • RAM is orders of magnitude faster than persistent disk storage.
  • Target a 95%+ Cache Hit Ratio for optimal performance and database protection.
  • The Pareto principle allows caching a small fraction (20%) of data to satisfy the vast majority (80%) of traffic.
  • Caching saves cloud infrastructure costs and absorbs sudden traffic spikes.

Topic Knowledge Check

Exercise 1 of 1 • Test your architectural comprehension.

Exercise 1 of 10 answered
1

What is the primary indicator of a successful and effective caching layer?

Rate This Architecture ChapterFeedback & Rating

How clear and actionable was this distributed systems breakdown?

Interactive Engineering Workbenches: