Notification System Fan-Out Lab (Interactive)
Survive a Twilio outage and a viral like burst while frequency caps and queue isolation protect OTPs. Simulate one minute of notification traffic through shared versus per-channel queues, measuring dedup collapse, spam-guard drops, DLQ depth, and push p95 latency.
Multi-Channel Notification Fan-Out
Drive preference guards, dedup, and channel-queue isolation through an SMS provider outage.
How It Works Under the Hood
A notification engine sits between event streams and unreliable third-party gateways with wildly different speeds: APNs responds in 80 ms, SMS takes seconds under 10DLC throttles and can hard-outage. If all channels share one worker queue, retrying SMS consumers head-of-line block authentication pushes, so channel isolation is a correctness requirement. On the user-protection side, Redis sliding-window frequency caps and 60-second content-hash deduplication fold a fifty-like viral burst into one summary instead of fifty wakes at 2 AM.
Core Architectural Principles
- Per-channel queues give each gateway independent retry, rate, and DLQ lifecycles.
- Content-hash dedup with TTL collapses repeat notifications into a single summary.
- APNs 410 Gone feedback prunes device tokens after app uninstalls.
Start notification designs by separating ingestion from dispatch through Kafka, then segregate push, SMS, and email into isolated worker pools with circuit breakers and exponential backoff with jitter. Add preference checks, quiet hours, and dedup before fan-out, and close with token invalidation loops. Quantify the shared-queue failure: one provider outage delaying 2FA OTPs is an availability incident, not a nuisance.
Isolated queues and guards maximize deliverability of critical alerts at the cost of pipeline duplication and tuning.