Local (In-Memory) vs Distributed Cache Lab (Interactive)
Add or remove an in-process L1 heap tier, move hit ratio, and watch Redis traffic, latency, and stale-fan-out react. Balance nanosecond local caching against shared Redis state: compute blended latency, network savings, and the pub/sub invalidation that reconciles them.
Two-Tier Cache: L1 Heap + L2 Redis Lab
Nanosecond local caches kill Redis traffic — but fragment state across pods. Tune the tiers and the Pub/Sub fix.
Fleet & Traffic
Aggregate L1 heap = 100 pods × 512 MB = 50.0 GB — isolated copies, lost on pod restart.
With Pub/Sub broadcast, a mutation publishes `PUBLISH invalidations user:42` and every pod evicts its L1 copy within milliseconds — the two-tier design Uber runs for geofence polygons.
How It Works Under the Hood
A local cache (Caffeine, Go sync.Map) lives in the application process: reads cost ~50–100 nanoseconds with zero network, but state fragments across autoscaled pods, competes with the app heap, and vanishes on restart. A distributed cache (Redis) gives one shared view across a thousand pods at ~0.5–2ms dominated by TCP and serialization. Uber-style systems stack them: L1 answers hot ultra-stable keys (geofence polygons, configs), L2 holds sessions and carts. The glue is Redis Pub/Sub broadcast — PUBLISH invalidations user:42 makes every pod evict its local copy within milliseconds, collapsing the staleness window from a full L1 TTL.
Core Architectural Principles
- Blended latency = L1-hit × 100ns + miss × (Redis 1ms + residual 25ms DB), so L1 ratio drives p50 directly.
- L1 absorbs its hit share of traffic as Redis network calls — typically 80–95% of socket load disappears.
- Local copies require pub/sub broadcast eviction after mutations, or one write leaves N pods serving stale data.
Propose L1+L2 tiering for 100k+ QPS read paths — ad bidding, feeds, geofences — and quantify: "L1 serves 80% of reads in 100ns and cuts Redis traffic accordingly." Then prove you know the consistency cost: each pod holds isolated copies, so mutations must PUBLISH invalidations to a Redis channel that all instances subscribe to, bounding staleness to milliseconds instead of the L1 TTL.
Nanosecond locality and massive network savings versus fragmented per-pod state that only broadcast invalidation can keep honest.