LinkedIn Kafka Log Topology Lab (Interactive)
Kill the O(N^2) ESB point-to-point web and price Kafka's one-log-fits-all topology. Compare N systems wired point-to-point against append-only Kafka topics every consumer reads independently, with zero-copy sendfile delivery and partitioned throughput across brokers.
LinkedIn: Point-to-Point vs Kafka Log
Count the O(N²) ETL spaghetti 2010 LinkedIn drowned in, then stream the same traffic through one append-only log with zero-copy brokers.
How It Works Under the Hood
LinkedIn built Kafka because its data integration was a quadratic tangle: every pipeline — activity feeds, metrics, Change Data Capture — needed its own connector to every other system. Kafka inverts this into a distributed commit log: producers append once, and consumer groups read at their own offsets, turning N x M connectors into N + M subscribers. Sequential disk writes plus page-cache residency and zero-copy sendfile sustain trillions of events per day, and the derived-data store Venice materializes the same streams into serving systems.
Core Architectural Principles
- Connector economy: point-to-point integration costs N*(N-1)/2 links while a shared log costs N producers plus M consumers.
- Zero-copy path: sendfile moves segments page cache to NIC without user-space copies, cutting CPU per byte delivered.
- Partition throughput: broker count scales with aggregate event rate and fan-out depth, since every consumer group rereads the full log.
When you propose Kafka, justify it structurally: "This eliminates dual-write fan-out because search, analytics, and feeds are just consumers on one log with independent offsets." Mention the LinkedIn origin and the 7-trillion-events-per-day number, and note the log becomes the source of truth that derived stores like Venice rebuild from. Interviewers score that as platform thinking.
A universal event bus decouples teams and enables replay, but every consumer must handle at-least-once delivery, schema evolution, and log retention costs.