Home/Labs/Trace Sampling Workbench
All 280 Labs
INTERACTIVE LAB🧭

Distributed Tracing Sampler Lab (Interactive)

Flip head vs tail sampling on 500 real traces and see which errors survive. Run a deterministic request cohort through head-based probabilistic sampling or OTel Collector tail-based rules, then inspect the surviving span waterfall.

Head vs Tail Sampling on a 500-Request Cohort

A biased coin at the gateway vs an OTel Collector that waits for the whole trace. Tune both, then inspect the kept span waterfall.

Traces kept5/500 (1.0%)
Error traces captured0/21
Slow (>p99) captured0/49
Storage cost / day @500k rps$2,990

Head sampling decided 99% of requests invisible BEFORE execution — 21 of 21 crashing requests left zero tracing data (and zero exemplar links from the Prometheus histogram).

traceparent: 00-00000000000000000000001274621895-0000004033325725-01 · OK
kept #1/5
[API Gateway (root)] GET /api/v1/orders/8842130 ms
[Auth Service] POST /verify-token12 ms
[Order Service] GET /orders/8842100 ms
[PostgreSQL] SELECT * FROM orders WHERE id=884210 ms
[Payment API (Stripe)] POST /charges/verify63 ms

How It Works Under the Hood

A trace is a DAG of timed spans stitched across services by the W3C traceparent header, but storing all 500k spans-per-second traces costs real money — petabytes daily. Head sampling flips a biased coin at the root span and inherits it downstream: cheap, but blind to the 96% of rare crashes it never recorded. Tail sampling buffers every TraceID in stateful collectors for 10-30 seconds, then keeps 100% of error or slow traces and thins boring fast OKs to a token rate, capturing every outage while cutting storage 95%. This lab shows captured-error counts, keep rates, and daily dollar cost side by side.

Core Architectural Principles

  • traceparent carries version, 16-byte TraceID, 8-byte parent SpanID, and sampled flags across HTTP, gRPC, and Kafka record headers.
  • Head sampling decides before execution and cannot retroactively rescue a dropped crashing trace.
  • Tail sampling rules keep 100% of ERROR-status or >p99-latency traces and 0.1% of fast successes.
Interview Round Script

Explain the traceparent header format byte by byte, then contrast sampling strategies with numbers: head at 1% loses ~99% of rare errors, while tail buffering keeps all error and latency-outlier traces at similar storage cost. Mention exemplars linking histogram buckets to TraceIDs and trace_id injected into log MDC to tie the three pillars together.

Key Trade-Offs

Near-zero-overhead head sampling that is blind to rare failures versus stateful collector buffering that guarantees error capture.

Related Curriculum Chapter

Distributed Tracing: Spans, Traces, & OpenTelemetry (OTel)

Read Full Chapter Blueprint

Explore More Interactive Labs

View All 280 Labs