Datadog: Realtime Distributed Tracing, High-Cardinality Metrics & APM Ingestion
How Datadog ingests tens of trillions of events daily across millions of client agents, aggregating high-cardinality metrics in memory and persisting logs in partitioned columnar stores.
Datadog uses edge daemon agents to aggregate local metrics in 10-second windows before sending compressed batches to regional ingestion gateways, protecting backend storage from unbuffered network bursts.
Edge Metric Aggregation & Adaptive Tail Sampling
Tens of trillions of telemetry events ingested dailyRecording 100% of distributed microservice traces at millions of requests per second would saturate network bandwidth and storage budgets.
Adaptive tail sampling keeps traces in memory buffers until request completion; only traces exhibiting HTTP 5xx errors or > 99th percentile latency are forwarded to persistent storage.
Requires memory overhead on sampling nodes, but slashes storage costs by 90% while capturing 100% of production anomalies.
Differentiate head sampling (sampling at the edge before request completes) from tail sampling (deciding whether to retain after observing full execution latency/status).
Memory Hierarchy: CPU, RAM, Disk, Cache
Explore the trade-offs between speed, cost, and capacity across CPU registers, L1/L2/L3 caches, Main Memory (DRAM), NVMe SSDs, and Magnetic Hard Drives.
Multi-Primary (Multi-Master) Replication
Scale write capacity across geographic regions: Active-Active topologies, write conflicts, Last-Write-Wins (LWW), Vector Clocks, and CRDTs (Conflict-Free Replicated Data Types).
Vector Databases & Similarity Search (Pinecone, Qdrant, pgvector)
Power modern AI & RAG systems: High-dimensional vector embeddings, Cosine Similarity, Approximate Nearest Neighbors (ANN), HNSW graph indexes, and hybrid search.
Polyglot Persistence: Composing the Modern Data Tier
The master architecture capstone: Synthesizing PostgreSQL (ACID), Redis (Cache), Cassandra (Time-series), Elasticsearch (Search), and S3 (Blob) into a cohesive distributed system.
Vector Clocks & Lamport Timestamps
Order events without synchronized physical clocks: Logical clocks, partial ordering, causality tracking, and concurrent write conflict detection.
Kafka Deep Dive: Topics, Partitions, Brokers, & Consumer Groups
Master Kafka internals: Topic partitioning, segment storage, leader/follower replication, ISR (In-Sync Replicas), acks=all durability, and consumer group rebalancing.
Kafka Offsets & Event Replay
Leverage log immutability: Commit offset internals (`__consumer_offsets`), auto vs manual commits, time-travel offset rewinds, and disaster recovery replay.
API Gateway Responsibilities & Edge Architecture
Centralize perimeter concerns: TLS termination, authentication, rate limiting, request routing, header sanitization, and Envoy/Kong.
Service Mesh: Sidecars, Envoy, & Istio
Offload networking from application code: Control plane vs Data plane, mTLS zero-trust encryption, traffic shifting, and distributed telemetry.
Datadog Engineering Team • 2023
Ready to Practice Datadog-Style Systems?
Start with foundational networking, compute, and storage, and build up to complex distributed consensus.