Why Caching Matters Lab (Interactive)
Slide QPS, hit ratio, and DB latency to watch database load, node counts, and monthly cost collapse. Model the latency and cost multiplier of caching: how each point of hit ratio changes database offload, RAM sizing, and cluster spend.
Cache Hit Ratio Economics Lab
See why a 99% hit ratio is 10× better than 90%: every point of misses is paid for in database cores.
Workload Inputs
95% is the industry target zone. Dropping to 90% would push 10.0k QPS at the DB instead of 5.0k — often forcing a cluster upgrade.
How It Works Under the Hood
Caching works because of physics: DRAM answers in ~100 nanoseconds, an SQL query in 5–50 milliseconds — a 50,000× gap. The Pareto principle says ~80% of read traffic targets ~20% of the data, so a modest RAM working set intercepts most queries. Hit ratio is the economic lever: at 100,000 QPS, 90% still sends 10,000 QPS to the database while 99% sends 1,000 — a 10× load reduction that defers expensive vertical scaling and absorbs flash-sale surges that would exhaust connection pools.
Core Architectural Principles
- DB load = QPS × (1 − hit ratio): misses, not traffic, size the database tier.
- 99% hit ratio is 10× better than 90% because residual load scales with the miss percentage.
- The hot working set (~20% of keys under 80/20) fits in cheap RAM and shields disk from surge traffic.
When introducing a cache, always state a target hit ratio (>95%) and derive DB offload out loud: "100k QPS at 99% leaves 1k QPS on PostgreSQL — a 10× reduction versus 90%." Justify cache RAM with the 80/20 working-set estimate and frame Redis as a cost multiplier preventing a five-figure database upgrade.
Sub-millisecond reads and massive cost savings come bundled with staleness, invalidation logic, and eviction management.