Streaming Backpressure Lab (Interactive)
Overdrive a slow sink and watch buffers OOM, credits throttle, or shedders drop data. Run a producer into a lagging consumer under none, pull-based, credit-based, and shed-during-congestion strategies, and compare OOMs, broker backlog, and effective throughput.
Backpressure: 10k eps vs a 1k eps Sink
Pick a flow-control strategy and watch whether the excess lands on disk, pauses the source, gets dropped — or kills the JVM.
Processed
0
0 eps sustained
Broker Backlog
—
safe on Kafka disk log
Shed / Dropped
0
1-in-N sampling loss
Stage Status
RUNNING
lag ≈ 0s
Downstream input buffers announced as credits; when they hit 0 the upstream operator halts and TCP’s zero-window even throttles the gateway — end-to-end flow control, latency instead of loss.
How It Works Under the Hood
When the producer outruns the consumer, somewhere must absorb the difference or everything dies. Unbounded buffering moves the failure to the sink’s heap: full memory then crash. Pull-based flow control pushes the buffer upstream into the broker—consumers fetch at their own pace and lag becomes an observable number. Credit-based and Reactive Streams request(n) contracts make the receiver advertise capacity so the sender throttles precisely, while load shedding trades completeness for liveness by dropping low-value events during bursts.
Core Architectural Principles
- Ingress above sink rate fills the local buffer linearly; without flow control the outcome is OOM, not patience.
- Pull semantics relocate the backlog to the durable broker: consumer lag grows but nothing crashes or is lost.
- Credit / request(n) propagation caps in-flight records to advertised capacity; shedding drops a measured share instead.
Say "backpressure is a signal from consumer to producer, and I choose where the queue forms." Compare broker-side lag (durable, observable) against in-memory buffering (fragile), cite Reactive Streams request(n) or TCP zero-window as the honest signal, and reserve shedding for telemetry where partial data beats a stalled pipeline.
Throttling producers keeps pipelines alive but raises end-to-end latency, while shedding protects latency at the cost of completeness.