Event-Driven Autoscaling Trigger Lab (Interactive)
Advance a worker queue minute by minute and see CPU scaling miss surges that KEDA backlog-depth scaling catches instantly. Contrast lagging CPU/memory triggers with leading queue-depth and consumer-lag triggers, capped by Kafka partition count.
Lagging CPU vs Leading Queue-Lag Autoscaling (KEDA)
Step the clock one minute at a time and compare scaling workers on CPU (lagging) vs backlog depth (leading).
Queue backlog
0
Avg worker CPU
0%
Worker pods
5
min 0 · SLA limit
ok ✓ (20,000)
Last 1 minutes (backlog ▍ / pods ▂)
Proactive: desired = ceil(backlog / 200) capped at 64 partitions (one consumer per partition — extra pods would idle). At empty queue KEDA scales to 0 and idle compute spend is $0.
How It Works Under the Hood
CPU is a lagging metric that betrays asynchronous workers: an I/O-bound fleet waiting on databases and third-party APIs holds about 25 percent utilization even as hundreds of thousands of messages pile up, so a 70 percent CPU trigger never fires and SLAs break. Leading indicators measure pending work before processing. KEDA feeds custom metrics to the HPA: for Kafka it sums LogEndOffset minus committed offset across partitions, and for SQS it divides visible message count by a target backlog per pod to compute replicas. Scaling is bounded by partition count, and empty queues scale to zero to erase idle spend.
Core Architectural Principles
- Target replicas = ceil(backlog / targetBacklogPerPod), capped at the number of Kafka partitions.
- I/O-bound workers stay under 30 percent CPU during a surge, so CPU triggers never fire.
- KEDA minReplicas 0 enables true scale-to-zero, freeing compute when the queue drains.
For any async worker fleet, recommend queue backlog or consumer lag over CPU, and explain why: I/O-bound workers keep CPU low while the backlog explodes and SLAs break. Give the KEDA formula ceil(backlog / target-per-pod) and note the hard ceiling: you cannot run more active consumers than Kafka partitions. Mentioning scale-to-zero for idle queues demonstrates cloud cost awareness interviewers look for.
Leading-metric autoscaling prevents backlog blowouts and enables scale-to-zero, but adds a KEDA operator and is capped by partition count.