Home/Labs/Kafka Log Topology
All 280 Labs
INTERACTIVE LAB📨

LinkedIn Kafka Log Topology Lab (Interactive)

Kill the O(N^2) ESB point-to-point web and price Kafka's one-log-fits-all topology. Compare N systems wired point-to-point against append-only Kafka topics every consumer reads independently, with zero-copy sendfile delivery and partitioned throughput across brokers.

LinkedIn: Point-to-Point vs Kafka Log

Count the O(N²) ETL spaghetti 2010 LinkedIn drowned in, then stream the same traffic through one append-only log with zero-copy brokers.

Point-to-point ETL connectors (O(N²))190
Publish-subscribe pipes (O(N))20
New connectors if 1 system is added+20 (p2p) vs +1 (log)
Ingest throughput960 MB/s → 2880 MB/s replicated
Consumer egress (read amplification)4800 MB/s
Brokers required (sequential log I/O)8
Memory copies per fetch1 (page cache → NIC via DMA)
Log topology: 20 pipes instead of 190 ETL connectors, 8 brokers streaming 4800 MB/s egress with 1 (page cache → NIC via DMA).

How It Works Under the Hood

LinkedIn built Kafka because its data integration was a quadratic tangle: every pipeline — activity feeds, metrics, Change Data Capture — needed its own connector to every other system. Kafka inverts this into a distributed commit log: producers append once, and consumer groups read at their own offsets, turning N x M connectors into N + M subscribers. Sequential disk writes plus page-cache residency and zero-copy sendfile sustain trillions of events per day, and the derived-data store Venice materializes the same streams into serving systems.

Core Architectural Principles

  • Connector economy: point-to-point integration costs N*(N-1)/2 links while a shared log costs N producers plus M consumers.
  • Zero-copy path: sendfile moves segments page cache to NIC without user-space copies, cutting CPU per byte delivered.
  • Partition throughput: broker count scales with aggregate event rate and fan-out depth, since every consumer group rereads the full log.
Interview Round Script

When you propose Kafka, justify it structurally: "This eliminates dual-write fan-out because search, analytics, and feeds are just consumers on one log with independent offsets." Mention the LinkedIn origin and the 7-trillion-events-per-day number, and note the log becomes the source of truth that derived stores like Venice rebuild from. Interviewers score that as platform thinking.

Key Trade-Offs

A universal event bus decouples teams and enables replay, but every consumer must handle at-least-once delivery, schema evolution, and log retention costs.

Related Curriculum Chapter

LinkedIn: The Origins of Apache Kafka & Universal Event Streams

Read Full Chapter Blueprint

Explore More Interactive Labs

View All 280 Labs