Competing Consumers & Visibility Timeout Lab (Interactive)
Rent messages with leases, kill workers mid-task, and watch expired visibility trigger redelivery. Feed a work queue a fleet of competing consumers, tune the visibility timeout, then OOM-kill a worker to see exactly-once processing’s dirty secret: at-least-once redelivery.
Competing Consumers & Visibility Timeouts
Workers dequeue messages, hold invisibility leases, and crash mid-flight — watch un-ACKed messages resurface.
Available
8
ApproxNumberOfMessages
ACK'd
0
DeleteMessage issued
Redeliveries
0
duplicate handlers required
Autoscaler Target
2
⌈depth × task-time ÷ 60s SLA⌉
QUEUE: worker fleet lease state
W1IDLE
ReceiveMessage (long poll 20s)
W2IDLE
ReceiveMessage (long poll 20s)
W3IDLE
ReceiveMessage (long poll 20s)
Tasks run 4–21s. Rule of thumb: timeout ≈ 3× P99. At 12s the timeout is SHORTER than long tasks → healthy workers see duplicate deliveries.
SQS hides a message for the visibility timeout the moment a worker polls it. An ACK (DeleteMessage) removes it forever; a crash leaves no ACK, so the lease expires and a surviving worker retries with ReceiveCount + 1. That is at-least-once delivery — handlers must be idempotent, and the autoscaler should read queue depth, not CPU.
How It Works Under the Hood
Competing consumers let a stateless worker pool scale horizontally: each receive() hands one message to one worker, so throughput multiplies with fleet size. Queue-based load leveler semantics come from the visibility timeout—the message is leased, not consumed. If the worker crashes or overruns the lease, the broker re-exposes it and another consumer picks it up, guaranteeing delivery but not single delivery. Setting the timeout shorter than task duration manufactures duplicate work; setting it huge hides crashed workers for minutes.
Core Architectural Principles
- Receive stamps an invisible-until deadline (a lease); the message is only deleted on explicit delete after success.
- Lease expiry silently redelivers in-flight work: duplicated side effects unless consumers are idempotent.
- Autoscaler sizing ⌈depth × avg task time / 60⌉ workers keeps the queue from becoming the latency.
Say "the broker guarantees delivery, not processing" and walk through the lease: visibility timeout, ack on success, redelivery on expiry. Point out that a too-short timeout is worse than a too-long one because it actively creates duplicates, and close by making handlers idempotent with a dedupe key so redelivery is harmless.
Competing consumers give linear horizontal scaling and automatic failover but force at-least-once semantics and idempotent handlers.