Observability & High-Cardinality TelemetryPRODUCTION RETROSPECTIVE

Datadog: Realtime Distributed Tracing, High-Cardinality Metrics & APM Ingestion

How Datadog ingests tens of trillions of events daily across millions of client agents, aggregating high-cardinality metrics in memory and persisting logs in partitioned columnar stores.

High-Level Architectural Overview

Datadog uses edge daemon agents to aggregate local metrics in 10-second windows before sending compressed batches to regional ingestion gateways, protecting backend storage from unbuffered network bursts.

Key Engineering Problems & Trade-Offs

Edge Metric Aggregation & Adaptive Tail Sampling

Tens of trillions of telemetry events ingested daily
The Scaling Problem

Recording 100% of distributed microservice traces at millions of requests per second would saturate network bandwidth and storage budgets.

Engineering Solution

Adaptive tail sampling keeps traces in memory buffers until request completion; only traces exhibiting HTTP 5xx errors or > 99th percentile latency are forwarded to persistent storage.

Architectural Trade-Offs

Requires memory overhead on sampling nodes, but slashes storage costs by 90% while capturing 100% of production anomalies.

How to Say This in an Interview

Differentiate head sampling (sampling at the edge before request completes) from tail sampling (deciding whether to retain after observing full execution latency/status).

Curriculum Topics Used in Datadog Architecture (27)
Full Syllabus
Phase 2#14

Memory Hierarchy: CPU, RAM, Disk, Cache

Explore the trade-offs between speed, cost, and capacity across CPU registers, L1/L2/L3 caches, Main Memory (DRAM), NVMe SSDs, and Magnetic Hard Drives.

9 min readRead Blueprint →
Phase 5#59

Multi-Primary (Multi-Master) Replication

Scale write capacity across geographic regions: Active-Active topologies, write conflicts, Last-Write-Wins (LWW), Vector Clocks, and CRDTs (Conflict-Free Replicated Data Types).

9 min readRead Blueprint →
Phase 5#73

Vector Databases & Similarity Search (Pinecone, Qdrant, pgvector)

Power modern AI & RAG systems: High-dimensional vector embeddings, Cosine Similarity, Approximate Nearest Neighbors (ANN), HNSW graph indexes, and hybrid search.

9 min readRead Blueprint →
Phase 5#74

Polyglot Persistence: Composing the Modern Data Tier

The master architecture capstone: Synthesizing PostgreSQL (ACID), Redis (Cache), Cassandra (Time-series), Elasticsearch (Search), and S3 (Blob) into a cohesive distributed system.

10 min readRead Blueprint →
Phase 6#81

Vector Clocks & Lamport Timestamps

Order events without synchronized physical clocks: Logical clocks, partial ordering, causality tracking, and concurrent write conflict detection.

9 min readRead Blueprint →
Phase 8#113

Kafka Deep Dive: Topics, Partitions, Brokers, & Consumer Groups

Master Kafka internals: Topic partitioning, segment storage, leader/follower replication, ISR (In-Sync Replicas), acks=all durability, and consumer group rebalancing.

10 min readRead Blueprint →
Phase 8#114

Kafka Offsets & Event Replay

Leverage log immutability: Commit offset internals (`__consumer_offsets`), auto vs manual commits, time-travel offset rewinds, and disaster recovery replay.

8 min readRead Blueprint →
Phase 9#135

API Gateway Responsibilities & Edge Architecture

Centralize perimeter concerns: TLS termination, authentication, rate limiting, request routing, header sanitization, and Envoy/Kong.

9 min readRead Blueprint →
Phase 10#142

Service Mesh: Sidecars, Envoy, & Istio

Offload networking from application code: Control plane vs Data plane, mTLS zero-trust encryption, traffic shifting, and distributed telemetry.

10 min readRead Blueprint →
Primary Technical Sources & Published Papers
Building Observability for Millions of Hosts

Datadog Engineering Team • 2023

Ready to Practice Datadog-Style Systems?

Start with foundational networking, compute, and storage, and build up to complex distributed consensus.

Start Free: Topic #1