Enterprise Graph & Realtime StreamingPRODUCTION RETROSPECTIVE

LinkedIn: Apache Kafka Origination, Espresso NoSQL & Distributed Social Graph

How LinkedIn invented Apache Kafka to handle trillions of daily metrics, built Espresso for distributed document storage, and queries 1B+ professional relationships.

High-Level Architectural Overview

LinkedIn connects members through a multi-tier architecture: the member connection graph is served in memory by custom graph engines, transactional profiles are stored in partitioned Espresso clusters, and all event changes are broadcast across global Kafka clusters.

Key Engineering Problems & Trade-Offs

Distributed Change Data Capture (Databus to Kafka)

Trillions of daily stream events processed across Kafka clusters
The Scaling Problem

Propagating database updates to downstream search indexes, feed algorithms, and analytics with zero data loss and sub-second delay.

Engineering Solution

Databus captures database transaction log changes at the storage engine level and streams them into Kafka topics without modifying application code.

Architectural Trade-Offs

Requires maintaining binary replication parsing compatibility with database engine versions, but guarantees near-instant read replica freshness.

How to Say This in an Interview

Explain why Change Data Capture (CDC) with Kafka beats application double-writing: double-writing causes silent data divergence on network failures.

Curriculum Topics Used in LinkedIn Architecture (1)
Full Syllabus
Primary Technical Sources & Published Papers
Kafka: A Distributed Messaging System for Log Processing

Kreps, Narkhede, Rao (ACM NetDB) • 2011

Ready to Practice LinkedIn-Style Systems?

Start with foundational networking, compute, and storage, and build up to complex distributed consensus.

Start Free: Topic #1