Two-Phase Commit Coordinator Blocking Lab (Interactive)
Broadcast prepare, collect votes, then kill the coordinator at the one moment that freezes everyone. Drive Prepare/Commit across three XA databases, inject a veto or a coordinator crash, and watch in-doubt cohorts pin their row locks until recovery.
Two-Phase Commit Blocking Lab
Drive Prepare/Vote then Commit/Abort across three XA cohorts — and kill the coordinator at the one moment that turns atomicity into a hostage situation.
Ready to initiate the distributed transaction over Order / Inventory / Payment.
- → Coordinator idle. Configure the failure injections on the right, then Execute 2PC.
Throughput sits at ~50–500 tx/sec precisely because every cohort’s locks survive a second network roundtrip. Spanner de-risks the crash by replicating the coordinator’s decision log through Paxos; microservices skip all of it and run Sagas instead — no locks, but compensation instead of rollback.
How It Works Under the Hood
Two-Phase Commit gives atomic all-or-nothing across heterogeneous resource managers: the coordinator broadcasts PREPARE, each cohort persists changes to its WAL and locks the touched rows to answer YES, and only unanimous YES lets Phase 2 broadcast COMMIT. The flaw is timing: if the coordinator dies after logging its decision but before broadcasting it, cohorts sit PREPARED forever — they cannot commit (another may have voted NO) nor abort (the decision may be COMMIT) — holding locks and blocking unrelated transactions. This lab reproduces that hostage window, counts locked rows and blocked time, and resurrects the coordinator from its persistent log.
Core Architectural Principles
- Prepare persists intent to WAL and takes exclusive locks; one VOTE_NO aborts everyone.
- Coordinator crash between logging and broadcasting leaves cohorts irreducibly in doubt.
- Recovery replays the durable decision log to unblock cohorts — locks held the entire outage.
Answer “why not 2PC in microservices?” with the blocking window and lock holding across a network round trip, then name alternatives and their costs: best-effort one-phase, Sagas with compensating transactions, or Spanner replicating the coordinator log through Paxos.
Cross-database atomicity is real, but every cohort’s locks become hostage to the coordinator’s uptime.