Paste Storage Sizing Lab (Interactive)
Size object storage, metadata, and read QPS, then compare lazy expiry to a sweeper. Turn pastes per day, average size, and TTL into steady-state terabytes while contrasting orphan-generating lazy expiry against a scheduled sweeper.
Pastebin: Metadata DB vs S3 Blob Storage Sizing
Split paste payloads from their index rows, pick a TTL, and watch steady-state object storage and cache-miss load respond.
› Lazy deletion: expired paste served as HTTP 410 on read, row tombstoned — S3 object stays billed until swept.
› Sweeper: cron batches DELETE FROM pastes WHERE expires_at < NOW() LIMIT 10000 against an indexed idx_expires_at during off-peak.
› Orphaned (expired-but-billed) objects: 0 TB — sweeper reclaims ~25% of aged volume
The core Pastebin insight is decoupling: a ~88-byte metadata row (7-char Base62 key, creator, TTL, click counter) lives in sharded Postgres/DynamoDB for fast listing, while the raw text blob lives in S3 keyed by that same Base62 ID. Hot pastes sit in Redis for sub-10ms reads; every cache miss is an S3 GET you pay for, so the TTL, hit rate, and delete strategy are cost levers, not afterthoughts.
How It Works Under the Hood
Pastebin is deceptively a storage-lifecycle problem. Object bytes accumulate as pastes/day x size x TTL-days, but the metadata database grows with every paste record even after the blob expires. The real fork is deletion policy: lazy expiry only reclaims a paste when it is next read — which for expired content is never — leaving orphaned blobs that silently double your footprint, while a sweeper job scans and purges on schedule. With reads running roughly ten times writes, the read path shapes the cache tier more than the write path.
Core Architectural Principles
- Steady-state storage = pastes/day x KB x TTL-days, converted to terabytes.
- Metadata rows persist (about 88 bytes of overhead each) even after the object blob is gone.
- Lazy purge leaves orphan terabytes; an active sweeper reclaims them on interval.
Call out the two-store split immediately — cheap object storage for blobs, a small relational or KV store for metadata and view counts. Then discuss expiry as the hard part: lazy deletion saves CPU but leaks storage, so a sweeper is the right answer. Note the read/write skew and why anonymous traffic makes caching and per-paste rate limiting essential.
Lazy TTL expiry is cheap to run but strands expired blobs, whereas an active sweeper reclaims space at the cost of scan load.