LinkedIn: Apache Kafka Origination, Espresso NoSQL & Distributed Social Graph
How LinkedIn invented Apache Kafka to handle trillions of daily metrics, built Espresso for distributed document storage, and queries 1B+ professional relationships.
LinkedIn connects members through a multi-tier architecture: the member connection graph is served in memory by custom graph engines, transactional profiles are stored in partitioned Espresso clusters, and all event changes are broadcast across global Kafka clusters.
Distributed Change Data Capture (Databus to Kafka)
Trillions of daily stream events processed across Kafka clustersPropagating database updates to downstream search indexes, feed algorithms, and analytics with zero data loss and sub-second delay.
Databus captures database transaction log changes at the storage engine level and streams them into Kafka topics without modifying application code.
Requires maintaining binary replication parsing compatibility with database engine versions, but guarantees near-instant read replica freshness.
Explain why Change Data Capture (CDC) with Kafka beats application double-writing: double-writing causes silent data divergence on network failures.
Kreps, Narkhede, Rao (ACM NetDB) • 2011
Ready to Practice LinkedIn-Style Systems?
Start with foundational networking, compute, and storage, and build up to complex distributed consensus.