Batch vs Stream Lambda/Kappa Lab (Interactive)
Set a freshness SLA and log volume, then change the business logic and price the backfill. Score Lambda’s batch-plus-speed dual stack against Kappa’s single replayable stream for your freshness SLA, event volume, and reprocessing cost when logic changes.
Batch vs Stream: Lambda → Kappa Evolution
Set the freshness SLA and event volume, then change the business logic and watch which architecture pays you back.
Lambda (batch + speed)
- codebases to keep in sync: 2
- exact merged freshness: 6.0 h (speed view is an approximation until the batch lands)
- logic-change backfill: 11.1 h @ 50k eps Spark rerun, twice
- SLA met? NO
Kappa (single log + stream)
- codebases: 1 — "a batch is just a bounded stream"
- freshness: 50 ms continuous Flink output
- logic-change backfill: 1.1 h via offset rewind at wire speed
- SLA met? YES (with watermark/state complexity)
Trigger a “business logic change” to price the reprocessing.
▸ verdict: Kappa meets freshness but the replay takes 1.1 h — pre-warm a second consumer group
Uber migrated marketplace analytics from Lambda's dual codebases to a Kafka + Flink Kappa pipeline: one code path computes surge multipliers live and, when the algorithm changes, rewinds to offset 0 and rebuilds history at wire speed. Keep batch for petabyte ad-hoc OLAP and training sets — the decision is freshness SLA versus state complexity, not fashion.
How It Works Under the Hood
Lambda guarantees exact results by reprocessing the whole batch layer periodically while a speed layer serves approximate fresher views—two codebases that drift until the batch lands. Kappa collapses the system to one stream processor over an immutable log: a "batch" is just a bounded stream replay from offset zero. The decision is arithmetic—if your freshness SLA exceeds the replay time of a single codebase, Kappa is strictly simpler; if petabyte batch window functions dominate, plain batch (or a hybrid) remains honest.
Core Architectural Principles
- Lambda backfill = volume / batch rerun rate, paid twice because SQL and the speed job must match.
- Kappa backfill = volume / replay rate via offset rewind—one codebase, then flip the serving alias.
- Freshness SLA versus backfill duration is the only scoring input that matters for the choice.
Cite Jay Kreps’ own retraction before recommending Lambda—dual codebases are maintenance debt—then decide by arithmetic: "if a Kappa replay of the full log finishes inside my SLA, I run one stream codebase." Reserve batch honestly for OLAP and ML training sets where hours of staleness are irrelevant.
Kappa offers one codebase and wire-speed replay but demands watermark sophistication, while Lambda guarantees exactness at double maintenance.