Home/Labs/Distributed Cron Leader Lab
All 280 Labs
INTERACTIVE LAB⏰

Distributed Cron & Leader Election Lab (Interactive)

Deploy cron on many replicas, turn off leader election, and duplicate a million invoices. Time-bucket millions of due tasks across scheduler replicas and workers: without leases every replica dispatches its own copy; crashes strand minutes or hand off in 10 seconds.

Cron at Scale: Leader Election & Worker Dispatch

Time-bucketed due tasks flow through 3 scheduler replicas into 5 worker pods. Duplicate a midnight invoice run by turning off leader election, then crash the leader.

Duplicate executions

0

Missed triggers

0

Queue backlog

0

Drain time

0.0 min

dispatch load vs worker capacity (150 tasks/min @ 30/worker)need ≥ 50 workers
DISPATCH TRACE (last 1 minutes):

min 00 · dispatched 0 · executed 0 · dup ×0 · missed 0 · backlog 0

leader holds lease; crash costs ~10s takeover delay, zero loss — PENDING rows survive in the store.

The golden rule: decouple the scheduler (lightweight timer engine, leader-elected via etcd/ZooKeeper lease) from the worker fleet (auto-scaling compute pulling from SQS/Kafka). Time-bucket the schedule store with an index on (trigger_time, status) so the master’s per-minute scan stays O(due) even at millions of alarms. For multi-day workflows (“welcome email → wait 3 days → nudge”), reach for Temporal/Cadence durable timers instead of hand-rolled cron.

How It Works Under the Hood

A single-server crontab is a silent SPOF, and deploying the app horizontally turns one nightly job into N concurrent charges. Production schedulers decouple the timer coordinator from the worker fleet: one leader, elected via an etcd lease, scans an index on (trigger_time, status), marks rows DISPATCHED, and enqueues payloads that an autoscaled worker pool drains. With durable PENDING rows, losing the leader costs a lease-TTL takeover delay—without election it costs duplicates or whole missed minutes.

Core Architectural Principles

  • No election: all replicas query the same time bucket, so duplicates = due tasks × (replicas − 1) every minute.
  • Lease-based election: standby promotes after the 10s TTL; PENDING rows survive, so dispatch is late but never lost.
  • Worker fleet capacity = pods × rate; dispatch beyond it builds a backlog whose drain time exposes under-provisioning.
Interview Round Script

State the golden rule first: "decouple the scheduler from the workers." Then handle correctness—leader election via etcd/ZooKeeper leases prevents duplicate dispatch, and time-bucketed indexed scans keep the master O(due). Close with Temporal or Cadence for multi-day durable timers instead of hand-rolled cron tables.

Key Trade-Offs

Leader election adds consensus infrastructure but converts duplicate-charge and missed-minute failures into bounded takeover delay.

Related Curriculum Chapter

Distributed Task Scheduling: Cron at Scale

Read Full Chapter Blueprint

Explore More Interactive Labs

View All 280 Labs