Home/Labs/2PC Blocking Lab
All 280 Labs
INTERACTIVE LAB⏳

Two-Phase Commit Coordinator Blocking Lab (Interactive)

Broadcast prepare, collect votes, then kill the coordinator at the one moment that freezes everyone. Drive Prepare/Commit across three XA databases, inject a veto or a coordinator crash, and watch in-doubt cohorts pin their row locks until recovery.

Two-Phase Commit Blocking Lab

Drive Prepare/Vote then Commit/Abort across three XA cohorts — and kill the coordinator at the one moment that turns atomicity into a hostage situation.

TRANSACTION COORDINATORIDLE

Ready to initiate the distributed transaction over Order / Inventory / Payment.

Order DB (PostgreSQL)
IDLE
Inventory DB (MySQL)
IDLE
Payment DB (Oracle XA)
IDLE
Failure injections (set before Phase 1)
Network roundtrips
0
In-doubt cohorts
0
Locked rows
0
Coordinator log
  • → Coordinator idle. Configure the failure injections on the right, then Execute 2PC.

Throughput sits at ~50–500 tx/sec precisely because every cohort’s locks survive a second network roundtrip. Spanner de-risks the crash by replicating the coordinator’s decision log through Paxos; microservices skip all of it and run Sagas instead — no locks, but compensation instead of rollback.

How It Works Under the Hood

Two-Phase Commit gives atomic all-or-nothing across heterogeneous resource managers: the coordinator broadcasts PREPARE, each cohort persists changes to its WAL and locks the touched rows to answer YES, and only unanimous YES lets Phase 2 broadcast COMMIT. The flaw is timing: if the coordinator dies after logging its decision but before broadcasting it, cohorts sit PREPARED forever — they cannot commit (another may have voted NO) nor abort (the decision may be COMMIT) — holding locks and blocking unrelated transactions. This lab reproduces that hostage window, counts locked rows and blocked time, and resurrects the coordinator from its persistent log.

Core Architectural Principles

  • Prepare persists intent to WAL and takes exclusive locks; one VOTE_NO aborts everyone.
  • Coordinator crash between logging and broadcasting leaves cohorts irreducibly in doubt.
  • Recovery replays the durable decision log to unblock cohorts — locks held the entire outage.
Interview Round Script

Answer “why not 2PC in microservices?” with the blocking window and lock holding across a network round trip, then name alternatives and their costs: best-effort one-phase, Sagas with compensating transactions, or Spanner replicating the coordinator log through Paxos.

Key Trade-Offs

Cross-database atomicity is real, but every cohort’s locks become hostage to the coordinator’s uptime.

Related Curriculum Chapter

Two-Phase Commit (2PC) & Its Limits

Read Full Chapter Blueprint

Explore More Interactive Labs

View All 280 Labs