Coordination Services ZooKeeper etcd Consul Lab (Interactive)
Load a workload onto ZK, etcd, or Consul and find the writes-per-second and value-size walls. Match leader election, config watches, session caching, discovery, or blob storage against three CP stores and compute latency, capacity use, and anti-pattern verdicts.
Coordination Store Sizing Bench
ZooKeeper, etcd, and Consul are small, quorum-fsync-ed, linearizable KV stores — brilliant for locks and evil for throughput. Load a workload onto one and see where it breaks.
- Write latency
- ~3 ms
- Ceiling utilisation
- 1%
- Value-size use
- 0%
- Verdict
- GOOD FIT
ZooKeeper on “Leader election / lock” — rare writes, strong guarantees
- No red flags: writes stay under quorum-fsync capacity and values fit the jumbo limit. one-shot: re-arm after every fire
Watch semantics differ more than APIs: ZooKeeper watches fire once and must be re-armed (a thundering re-arm after failover), etcd revision streams and Consul blocking queries re-issue cheaply. All three stop writing during leader election — which is why “who is the leader” belongs here and “what is the cart total” does not.
How It Works Under the Hood
ZooKeeper, etcd, and Consul are deliberately small linearizable key-value stores: every write is a quorum fsync, the dataset is bounded (tens of GB), and values face jumbo limits around 1–1.5 MB. Used as coordination primitives — locks, leader election, config propagation, membership — they are superb; used as general databases they cap throughput and couple availability to the quorum. Watch semantics also differ materially: ZooKeeper fires one-shot watches that must be re-armed, etcd streams continuous revisions, Consul re-issues blocking queries. This lab scores your workload pairing and explains every wall it finds.
Core Architectural Principles
- Quorum-fsync write ceiling (4k–12k/sec by service) versus session-cache write demand.
- Value-size caps expose coordination stores as a poor blob or media backend.
- One-shot versus streaming watch styles decide re-arm storms and failover overhead.
Bound the blast radius in design reviews: “we store only coordination state in etcd — leaders, locks, config — with writes under a few thousand per second; anything high-rate or large goes to Cassandra or S3.” Mention CP-store unavailability during leader elections as the reason not to put user reads there.
Linearizable coordination beats any hand-rolled alternative, but its small, fsync-bound scope is non-negotiable.