Gorilla TSDB Compression Lab (Interactive)
Shrink points 16 B to 1.37 B with Gorilla, then blow up the index with cardinality. Model time-series ingestion bytes with delta-of-delta and XOR compression and expose cardinality-driven RAM out-of-memory.
Prometheus/Datadog TSDB: Gorilla Compression & Cardinality Blow-ups
Compression shrinks bytes-per-point; unbounded labels explode the series index — the two forces that decide TSDB cost.
samples every 10s: t = [1000, 1010, 1020, 1030] → deltas [10,10,10] → delta-of-delta [0,0,0]
each zero DoD = 1 bit (leading 0); values via XOR of meaningful bits only.
Gorilla exploits two realities of telemetry: timestamps arrive at fixed intervals (so the delta-of-delta is almost always a single 0 bit) and consecutive gauge values barely move (so XOR leaves only a handful of meaningful bits) — collapsing 16 B to 1.37 B and cutting the daily write from 13.8 TB to 1.2 TB. Compression fixes the throughput problem; it does nothing for the cardinality problem, because every unique label combination is a separate series that must live in the RAM index regardless of how small its datapoints get. That is why the golden rule — never label metrics with unbounded identifiers — matters more than any encoding trick.
How It Works Under the Hood
Metrics pipelines face two independent pressures: throughput volume and index explosion. Gorilla compression exploits that timestamps tick at fixed intervals, so the delta-of-delta is almost always a single zero bit, and consecutive values barely change, so XOR leaves few meaningful bits — collapsing 16-byte points to about 1.37 and cutting 10M points/sec from roughly 13.8 TB/day to a tenth. But compression does nothing for cardinality: each unique label combination is a separate series resident in the RAM index, so labeling a metric with user_id across 10M users allocates 10M series and OOMs the node.
Core Architectural Principles
- Gorilla: delta-of-delta timestamps, one bit when zero, plus XOR-compressed floats turns 16 B into 1.37 B.
- Daily ingest = points/sec x bytes/point x 86,400, so 13.8 TB raw versus about 1.18 TB at 1.37 B.
- Active series = cardinality x metrics; past roughly 5M series the RAM index OOMs the node.
Separate the metrics TSDB path from the log and trace indexing path, then explain Gorilla — delta-of-delta plus XOR — concretely to show low-level fluency. Raise cardinality as the killer constraint and commit to never labeling with unbounded identifiers, with schema linting as the guardrail. Mention downsampling and rollups for long retention and the 13-month cost framing.
Gorilla compression slashes metric storage by about 90% but cannot tame cardinality, since every unique series still occupies RAM index space.