Batching & Bulk Operations Lab (Interactive)
Group 1,000 events into one TCP frame and one fsync to trade milliseconds of latency for hundredfold throughput. Compare single-row loops, multi-row INSERTs, and PostgreSQL COPY streams while tuning batch size, event rate, and Kafka linger.ms.
Batch Amortization Throughput Lab
Pay the fixed syscall / TCP-frame / fsync tax once per batch, not once per row.
How It Works Under the Hood
Every operation carries a non-negotiable fixed tax: 80+ bytes of Ethernet/IP/TCP/TLS framing, a user-to-kernel context switch, SQL parse, lock allocation, and a WAL fsync. Paying it per row puts 100k inserts at 145 seconds; paying it once per 1,000-row batch drops it to 3.2 seconds, and the binary COPY stream — no parser, no planner, one commit — lands at 0.75 seconds. Kafka producers formalize the trade as linger.ms and batch.size, and batched compression exploits cross-record schema redundancy for 4-6x ratios.
Core Architectural Principles
- Amortized cost per row = fixed overhead ÷ batch size + per-row marginal cost, so bigger batches flatten the curve.
- Dual-trigger micro-batching flushes on size OR timer so low-volume streams never sit in the buffer indefinitely.
- Batched LZ4/Zstd compression ratios multiply because repeated JSON keys and headers compress across the batch.
Asked how to ingest 50k rows/sec, answer in tiers: multi-row INSERT first, PostgreSQL COPY or Kafka batching for order-of-magnitude gains, and cite ClickHouse's 1k-100k row batch mandate. Mention SIGTERM buffer draining — losing in-flight batches on deploy is the follow-up question that separates seniors.
Batching trades 5-50 ms of added buffering latency and buffer-memory management risk for 10x-100x throughput, and is the wrong call for sub-millisecond trading paths.