Atomic Rate Limiter Lab (Interactive)
Race concurrent requests against an atomic Lua token bucket vs a GET-then-INCR check. Model refill rate, burst capacity, and concurrent request waves to expose the over-admission bug in a non-atomic limiter.
Redis Lua Token Bucket vs Check-then-Set Race
Drive one-second traffic waves at a rate limiter and watch atomic Lua scripts close the TOCTOU gap that naive read-write allows.
Send a wave: with the race mode, concurrent readers all compare the stale Redis counter against the limit before any decrement lands.
Token Bucket wins as the industry standard because it tolerates bursts (bucket depth) while enforcing a sustained rate (refill). In Redis Cluster the key rate_limit:{user_id} is sharded across 16,384 hash slots; a Lua script executes single-threaded per shard, making refill + compare + decrement one atomic step. Under Redis outage, gateways fail-open with a local in-memory fallback rather than taking the API down with the limiter.
How It Works Under the Hood
A token bucket is trivial to describe and easy to get wrong in a distributed setting. If a limiter does GET tokens and then INCR separately, concurrent requests read the same stale count and every one of them passes — the bucket over-admits by roughly the concurrency level. Executing the read-decide-write as a single atomic Lua script removes that window. Sweep burst, refill, and a simultaneous crowd to see exactly how many extra requests leak through the racy path and what Retry-After you owe them.
Core Architectural Principles
- Fair path admits up to the token budget; the GET-INCR race over-admits by (crowd - 1) leaked tokens.
- Bucket refills at rate x elapsed and is capped at burst capacity.
- Denied requests receive a Retry-After derived from the time until the next token lands.
The whole question is atomicity: narrate the check-then-act race in GET then INCR and fix it with a single Lua script or INCR-with-expire. Quantify over-admission under a concurrent wave to show you understand why it passes unit tests yet fails at scale. Mention Redis as a shared source of truth and why clock-skewed local counters drift.
A single Redis limiter is atomic and precise but adds a round trip per request; local token buckets are fast yet undercount across processes.