Home/Labs/Batch Amortization Lab
All 280 Labs
INTERACTIVE LAB📦

Batching & Bulk Operations Lab (Interactive)

Group 1,000 events into one TCP frame and one fsync to trade milliseconds of latency for hundredfold throughput. Compare single-row loops, multi-row INSERTs, and PostgreSQL COPY streams while tuning batch size, event rate, and Kafka linger.ms.

Batch Amortization Throughput Lab

Pay the fixed syscall / TCP-frame / fsync tax once per batch, not once per row.

Ingestion rate32,258 rows/s
100k rows in 3.1 s
Speedup vs loop43.9×
one parse + commit per batch
Added latency31.1 ms
wait ≤ min(fill 0.1 ms, linger) + exec
Protocol overhead0.3%
batch compresses 4.0× better than singles
20 batches/sec · 1,000 rows each · in-memory buffer 97.7 KB · drain on SIGTERM for zero lost events
Comfortably absorbing 20,000 events/s with 1 writer thread(s).

How It Works Under the Hood

Every operation carries a non-negotiable fixed tax: 80+ bytes of Ethernet/IP/TCP/TLS framing, a user-to-kernel context switch, SQL parse, lock allocation, and a WAL fsync. Paying it per row puts 100k inserts at 145 seconds; paying it once per 1,000-row batch drops it to 3.2 seconds, and the binary COPY stream — no parser, no planner, one commit — lands at 0.75 seconds. Kafka producers formalize the trade as linger.ms and batch.size, and batched compression exploits cross-record schema redundancy for 4-6x ratios.

Core Architectural Principles

  • Amortized cost per row = fixed overhead ÷ batch size + per-row marginal cost, so bigger batches flatten the curve.
  • Dual-trigger micro-batching flushes on size OR timer so low-volume streams never sit in the buffer indefinitely.
  • Batched LZ4/Zstd compression ratios multiply because repeated JSON keys and headers compress across the batch.
Interview Round Script

Asked how to ingest 50k rows/sec, answer in tiers: multi-row INSERT first, PostgreSQL COPY or Kafka batching for order-of-magnitude gains, and cite ClickHouse's 1k-100k row batch mandate. Mention SIGTERM buffer draining — losing in-flight batches on deploy is the follow-up question that separates seniors.

Key Trade-Offs

Batching trades 5-50 ms of added buffering latency and buffer-memory management risk for 10x-100x throughput, and is the wrong call for sub-millisecond trading paths.

Related Curriculum Chapter

Batching & Bulk Operations: Amortizing System Overhead

Read Full Chapter Blueprint

Explore More Interactive Labs

View All 280 Labs