Backup & PITR (RPO vs RTO) Lab (Interactive)
A DROP TABLE strikes at 20:00; pick the backup regime and compute lost minutes, replay time, and storage cost. Quantify Full vs Differential vs WAL-archived PITR strategies as RPO data-loss windows, RTO restore clocks, and dollar tradeoffs.
Backup Strategy & PITR Recovery Planner
An engineer runs DROP TABLE at 20:00. Choose the backup regime and compute RPO, RTO, and cost.
RPO (data loss)
1 min
0.00 GB of commits gone
RTO (downtime)
89 min
download + WAL replay
Weekly storage
541 GB
S3 snapshots + archives
Recovery granularity
any second
Without WAL archiving, your RPO equals the time between the last full/differential snapshot and the disaster — restoring a 23:59 crash from the midnight backup loses 1200 minutes of orders. Continuous WAL streaming to S3 pins RPO to the 1-minute archive gap at the cost of replay time during restore, which is exactly why a warm replica promoted via failover (RTO ~30s) beats any cold restore.
How It Works Under the Hood
Disaster recovery is defined by two numbers: RPO, the maximum committed data you can afford to lose, and RTO, how long service may stay offline. Nightly full backups pin RPO to nearly a day, differentials shorten storage but not the loss window, and continuous WAL archiving to object storage shrinks RPO to the archive interval while enabling Point-in-Time Recovery, replaying every log record up to one second before the errant DROP TABLE.
Core Architectural Principles
- RPO maps directly to backup/WAL archive frequency; RTO adds snapshot download plus WAL replay time.
- PITR reconstructs any timestamp inside the retention window, not just the last snapshot.
- Warm standby promotion delivers seconds-level RTO that no cold restore can match.
State numeric targets in every DR answer: "RPO under one minute via continuous WAL archiving to S3, RTO under thirty seconds via automated replica promotion." Explain why a midnight snapshot cannot save you from a 14:32 DROP TABLE, and contrast logical pg_dump restores with physical base-backup plus replay for terabyte databases.
Millisecond-precise recovery and seconds-level RPO demand WAL pipelines and warm standbys that pure snapshots never need.