Memory Hierarchy Lab (Interactive)
Sweep cache tiers with array vs linked-list access and NUMA penalties. Trace memory accesses through L1/L2/L3 and DRAM, watching spatial locality and working-set size decide where every hit lands.
Memory Hierarchy & Locality Simulator
Trace memory accesses through L1/L2/L3 → DRAM and watch locality, cache capacity, and NUMA change the physics.
Hierarchy — where misses are served
Spatial locality: each 64-byte cache line load covers 8 consecutive elements, so only 1 in 8 accesses misses the cache.
How It Works Under the Hood
CPUs execute in fractions of a nanosecond while DRAM costs ~100ns, so hardware stacks SRAM caches between them. Data moves in 64-byte cache lines: contiguous arrays let spatial locality and prefetching amortize misses to one per eight accesses, while pointer chasing pays full memory latency every time. Working-set size decides which tier serves misses, and NUMA remote sockets double DRAM latency.
Core Architectural Principles
- Cache lines are 64 bytes: an array scan takes 1 miss per 8 element accesses, a linked list misses on every dereference.
- The deepest tier that fits the working set becomes the miss-servicing latency floor.
- NUMA remote memory access traverses UPI/Infinity Fabric, roughly doubling DRAM latency.
When justifying data layout or index choices, quantify the gap: "array traversal is ~50x faster than pointer chasing because 64-byte cache lines give us spatial locality." Mention false sharing and NUMA pinning (numactl --membind) for databases handling 100k+ QPS to signal real systems depth.
Faster memory costs orders of magnitude more per byte and is volatile, so hot sets go to DRAM and cold data tiers to NVMe and HDD.