RED and USE Metrics Lab (Interactive)
Add forbidden labels until Prometheus OOMs, and watch averages hide tail latency. Switch between the RED service method and USE resource method, multiply label cardinality into a RAM explosion, and compare mean against p95/p99 percentiles.
RED vs USE Telemetry & Cardinality Explosion Lab
Pick the method, add forbidden high-cardinality labels to see Prometheus RAM explode, and compare mean vs percentile latency.
Requests/sec — the traffic the customer generates.
sum(rate(http_requests_total[5m])) by (service, route)Failed request ratio, not raw counts.
sum(rate(http_requests_total{status=~'5..'}[5m])) / sum(rate(http_requests_total[5m]))p99 latency distribution — never the mean.
histogram_quantile(0.99, sum(rate(http_request_duration_seconds_bucket[5m])) by (le))Panel 1: Time-series cardinality
Healthy: only low-cardinality enums (<50 values each) label this metric.
Panel 2: Why averages hide tail latency
p99 is 0.1× the mean. Tail and mean agree right now — push the slow-request sliders up to reproduce the classic masked outage.
Rule of thumb: RED for every user-facing service, USE for every finite resource (hosts, pools, queues) — saturation is the earliest warning.
How It Works Under the Hood
RED (Rate, Errors, Duration) describes every request-driven service from the customer perspective; USE (Utilization, Saturation, Errors) describes every finite resource, with saturation as the earliest pre-outage warning. But TSDB health hinges on label hygiene: each unique label-value combination allocates a live time series consuming 2-4 KB of RAM, so one accidental user_id label multiplies into 10^14 series and an unrecoverable WAL-replay crash loop. Percentiles complete the picture: 99 requests at 5 ms and one at 5 s average to a healthy-looking 55 ms while a user timed out.
Core Architectural Principles
- Cardinality is the product of label values; only low-cardinality enums (<50 values) belong in metric labels.
- Prometheus allocates kilobytes per active series, so unbounded labels trigger OOM crash loops on WAL replay.
- Latency SLIs require histogram_quantile over buckets (10/50/250ms/1s); arithmetic means mask catastrophic tails.
Lead telemetry design with the two-method split: RED for services, USE for resources. Then prove TSDB literacy by warning about high-cardinality label explosion and by insisting on percentile histograms instead of averages — quoting the 99-fast-plus-one-slow request example shows production instinct most candidates lack.
Rich, queryable label dimensions versus TSDB memory safety, and histogram bucket precision versus per-scrape cardinality cost.