Context Switching Overhead Lab (Interactive)
Oversubscribe cores with threads and watch switches, TLB flushes, and cache pollution eat CPU. Model direct register-save and indirect TLB/cache costs of preemption across OS threads, processes, and goroutines on a shared server.
Context Switch Overhead & Thrashing Simulator
Oversubscribe cores with threads and watch register saves, TLB flushes, and cold caches eat your CPU budget.
TLB stays warm — shared page tables
Mitigations from the topic: cap CPU-bound pools at Ncores, pin hot threads with taskset affinity, and replace thread-per-connection with epoll event loops — Nginx runs 1 worker per core with near-zero switches.
How It Works Under the Hood
Every preemption costs ~1-2 μs of kernel register saving, but the hidden tax is indirect: process switches reload CR3, flushing the TLB, and the next thread pollutes cold L1/L2 caches for 10-30 μs. Oversubscribe 8 cores with 10,000 threads and the scheduler devours most CPU cycles — thread thrashing. Goroutines switch in ~200ns because they stay in user space, and epoll event loops avoid switching almost entirely.
Core Architectural Principles
- Direct overhead is register save/restore; indirect overhead (TLB flush + cache warming) is 10-30x larger.
- Process switches reload page tables and flush the TLB; thread switches within a process keep it warm.
- Switch rate scales with cores and quantum, so wasted CPU = switches/s x cost per switch.
Diagnose latency stories with switch vocabulary: "we saw high involuntary switches in pidstat -w, meaning CPU oversubscription; we capped the pool at N_cores and pinned hot threads." Explain why goroutines switch in 200ns — no Ring 0 transition, ~14 registers, no page table touch.
Preemptive multitasking guarantees fairness and responsiveness, but unbounded thread counts collapse throughput through switching and cache pollution.