Write-Back (Write-Behind) Caching Lab (Interactive)
Dial write throughput and flush intervals, compute merged DB updates, then crash the node to size the loss window. Model asynchronous write buffering: sub-millisecond RAM acknowledges, batched coalesced flushes, and exactly how many acknowledged writes a crash eats.
Write-Back Coalescing & Loss Window Lab
Acknowledge in RAM, flush batches later — and measure exactly how much a crash eats.
Buffer Configuration
Lengthening the flush window drives IOPS toward zero but grows the loss window linearly — YouTube view counters happily eat a 5s loss; a payment ledger never will.
Durability mitigations when the loss window is unacceptable: Redis AOF appendfsync everysec, synchronous replica ack before client ACK, or a Kafka write-ahead log drained by the flusher.
How It Works Under the Hood
Write-Back acknowledges after RAM only, marking keys dirty, and an async flusher drains the queue periodically (say every 5 seconds) as consolidated bulk updates. YouTube’s view counter is the canonical math: 10,000 increments per second collapse into one UPDATE videos SET views = views + 50000, cutting database IOPS by 99.98%. The bill comes on crash: every acknowledged-but-unflushed write inside the flush window is permanently lost. Acceptable for counters and telemetry; a compliance violation for ledgers — mitigated with AOF everysec, synchronous replica acks, or a Kafka write-ahead log.
Core Architectural Principles
- Write coalescing: QPS × flush-interval increments per key merge into a single batched DB update.
- Data-loss window = write rate × flush interval, growing linearly with buffer depth.
- Durability mitigations: Redis AOF everysec, cross-AZ replica ack, or WAL to Kafka before acknowledging.
Recommend Write-Back whenever designing high-frequency counters — view counts, leaderboards, ad impressions — and recite the coalescing math: "10,000 in-memory increments per second flush to one SQL UPDATE every 5 seconds, a 99%+ IOPS reduction." Then explicitly name the data-loss window and justify it: losing three seconds of view counts is fine; losing three seconds of payments is not.
Extreme write throughput and minimal disk IOPS purchased with a bounded window of acknowledged-write loss on crash.