TOPIC #225Intermediate 13 min read

Batching Strategies: Static, Dynamic, Continuous

AI
AI & ML Editorial
Report an issue
Key takeawayCore Concept Summary

A GPU is like a bus: sending one request per trip wastes almost everything. Batching converts waiting into throughput — static (client assembles fixed batches), dynamic (server collects arrivals within a small delay budget), and continuous/in-flight (LLM decode pools where sequences join and leave every token step). This topic covers the economics, the padding-waste and tail-latency math each strategy trades, and the Triton/vLLM knobs that control them.

Three Ways to Fill a GPU

Static batching pushes waiting to the client; dynamic batching lets the server amortize arrivals over a bounded delay; continuous batching recycles the running set per decode step, eliminating head-of-line waits for variable-length sequences.

Three Ways to Fill a GPU
100%
Touchpad: Pinch to zoom • Drag to pan
Rendering visual architecture flowchart...

01.The Problem: A Bus That Leaves With One Passenger

A GPU is scary good at doing many things at once — and surprisingly bad at value-for-money when it does one.

Concrete numbers: push a single 64x512 matrix through a vision encoder and you may use ~5% of an H100 tensor-core throughput. Push 64 of them together as one batch and you approach saturation. Same chips, same weights, wildly different utilization.

Why does the batch version barely cost more?

  • The model's weight loads and kernel launches — fixed overhead per trip — are amortized across everyone in the batch.
  • Memory-bound elementwise work becomes compute-efficient once many rows share each memory fetch.

The rough law of the whole topic:

Insight

Batching multiplies throughput ~linearly until compute or HBM bandwidth saturates; per-item latency stays near flat — then grows with batch size once the GPU is full (queueing plus slower steps).

So the question is never "should I batch?" It is who waits, where, and how much is wasted while waiting. Three cost axes you manage:

  • Wait time (forming the batch) → lands on p50/p99 latency.
  • Padding waste (ragged shapes forced into one batch) → wasted FLOPs; mitigate with length bucketing.
  • Memory → KV cache and activations grow with batch; on LLMs, OOM decides max batch.

02.The Idea in Plain Words: Three Ways to Fill the Bus

All the jargon in this topic is three sentences:

  • Static batching: the client waits, assembles a fixed-size batch (say 32 requests), and ships it. The wait happens in the caller.
  • Dynamic batching: the server catches individual arriving requests, holds them in a queue for a bounded moment (a few milliseconds), and fires whatever has gathered as one GPU step. The wait happens in a controlled place.
  • Continuous (in-flight) batching: for LLMs that generate token by token — the batch is a pool that is recomposed after every single token step: finished sequences leave immediately, waiting ones join mid-flight.

So the taxonomy is just a progression of answers to one question

Insight

When does a request stop waiting and start computing?

Static says: when my batch is full. Dynamic says: when my delay budget expires. Continuous says: this very token step — nobody waits for the slowest passenger anymore.

Sections 4-5 build the picture; sections 6-8 give each strategy its knobs and math.

03.A Worked Example: The Cost of Riding Together

Padding waste, tiny numbers. Three text requests arrive with token lengths 10, 100, and 900. A batched GPU step must give every row the same shape, so they are padded to the longest: the batch computes as if all three were 900 tokens.

  • Useful work: 10 + 100 + 900 = 1,010 tokens.
  • Performed work: 900 x 3 = 2,700 tokens.
  • Efficiency: 1,010 / 2,700 ≈ 37% — nearly two-thirds of the FLOPs painted padding.

That is why length bucketing exists: batch the 10s with the 10s, the 900s with the 900s.

Idle-slot waste, the LLM version. Decode generates one token per sequence per step. Under static/dynamic batching, a 10-token answer and a 900-token answer join the same batch. The short one finishes at step 10 — and then holds its batch slot idle for 890 steps, blocked from release until the longest sibling finishes, while everyone in that batch waits on it to complete before the batch frees.

Continuous batching is the fix with a one-line mechanism: release the slot the instant the sequence emits its end token; admit a new request at that same step boundary. Idle slots → gone. Occupancy → near 100%.

04.The Analogy: City Transit, Shuttles, and Ride Pools

Carry this through: the GPU is a bus, and each request is a passenger.

  • Static batching = a chartered coach. You stand at the stop until 32 people show up, then everyone departs. At rush hour it is efficient; at 3 am you stare at the stop for ten minutes — or the driver leaves half-empty, wasting the trip anyway.
  • Dynamic batching = a shuttle with a rule: "departs every 3 minutes, or when full." Nobody waits more than the posted budget, and two passengers arriving 1 minute apart still share a ride. The server, not the passenger, manages the waiting — that is the whole advance.
  • Continuous batching = a real city bus. Passengers hop off at their stop and new ones hop on at every stop (every token step). The bus never idles waiting for the passenger who booked a 900-stop journey; the 10-stop rider is simply gone after stop 10.

Each analogy also names its failure mode: the chartered coach waits forever at an empty stop (low-QPS endpoint — the delay is pure loss); the shuttle's 3-minute rule is a dial you tune (too long hurts p99 at low traffic, too short underfills the bus); the city bus needs a smart driver and door schedule (a real scheduler: preemption, admission, KV memory).

05.Visual Intuition: Timelines of Three Buses

Same arrivals (r1..r5), three policies. Rows are GPU steps; letters are requests; . is a wasted beat.

code
STATIC (fixed batch of 4)      DYNAMIC (3 ms budget)      CONTINUOUS (per-step pool)
 r1--|                         r1--| step: [r1 r2]        r1 |r1.|           <- r1 ends early
 r2--|  wait, wait...          r2--| step: [r3 r4 r5]     r2 |r2r2|          <- r2 longer
 r3--|________________         r3--|  (no waiting r1!)    r3 |r3r3|
 r4--|  step: [r1 r2 r3 r4]    r4--|                      r4 |r4r4|
 r5................................  step: [r5 ...]       r5 joins mid-flight:
                                                        step |r1r2 + r5|
   r5 waits a whole batch cycle    steps fire on schedule  slots recycle every step

Read the top row of each: static trades r5's latency for a fuller bus; dynamic bounds the trade (never more than the delay budget); continuous removes the end-of-batch coupling entirely — the cost it removes is not the wait to start, it is the wait to finish.

06.Static and Dynamic Batching: The Server Learns to Wait Politely

Static batching: the client accumulates requests and sends shaped batches (batch_size=32). Optimal for offline scoring where latency is irrelevant; terrible online — clients either wait long or send underfilled batches.

Dynamic batching (Triton dynamic_batching, TF Serving batching parameters): individual requests hit the server; the scheduler holds them in a queue up to max_queue_delay_microseconds (e.g. 2-5 ms), forming batches from whoever arrived, respecting preferred_batch_size buckets. Two requests 1 ms apart share a GPU step — the wait is paid from inter-arrival slack, not added to a lonely request's latency, as long as arrivals keep the queue fed.

Key config semantics in Triton, one line each:

  • max_batch_size and preferred_batch_size buckets: what a full bus looks like.
  • max_queue_delay_microseconds: the explicit latency-for-throughput dial.
  • default_queue_policy (timeout/reject): so overload fails predictably instead of melting p99 — better a clean 503 than a mushy 2-second success.
  • Sequence/ragged batching for variable-length inputs: bucketing by length limits the padding waste from the worked example.
text— Triton dynamic batching config with a 3 ms delay budget
dynamic_batching {
  preferred_batch_size: [ 8, 32 ]
  max_queue_delay_microseconds: 3000
  default_queue_policy {
    timeout_action: TIMEOUT          # 503 rather than unbounded wait
    timeout_us: 10000
  }
}

07.Continuous (In-flight) Batching for LLMs

Transformer decode is autoregressive: one token per sequence per step — the setup for the idle-slot math in section 3. Continuous batching (Orca-style, adopted by vLLM, SGLang, TensorRT-LLM "in-flight batching", TGI) makes the batch a pool: after every token step, finished sequences release their slots (and KV pages) and queued requests are admitted mid-flight. The scheduler runs prefill/new-decode mixes each iteration (or chunked-prefills long prompts to protect decode latency).

Measured effect at NeurIPS 2023-era deployment: continuous batching plus paged KV raised serving throughput 2-4x at equal p99 versus per-request static batching, because SM occupancy stays near 100% and no batch slot idles on the longest sibling.

The three quantities schedulers actually trade:

08.Choosing and Tuning: Which Bus for Which City

Decision table for the design interview:

  1. Offline scoring / embeddings backfill: static large batches, spot GPUs, no latency dial — throughput only. (Charter the coach; nobody is watching the clock.)
  2. Classical online models (ranking, vision, ASR encoders): dynamic batching; set queue delay to the fraction of your latency budget left by everything else (network + feature fetch + post). Start at 2-4 ms, sweep, verify against the SLO under arrival-rate spikes.
  3. LLM serving: continuous batching is non-negotiable (it is built into vLLM/TRT-LLM); tune max_num_seqs, max_num_batched_tokens, long-prefill chunking, and preemption/recompute policy under memory pressure.

Always watch the batch telemetry — the dashboard that exposes silent waste:

  • Mean/median batch size, and % of batches hitting preferred_batch_size.
  • Queue depth at p99 load.
  • Padding ratio; for LLMs, token efficiency = generated tokens / computed tokens.

A fleet silently running at mean batch size 1.3 is leaving 2-3x of the GPU's capacity on the table — the cheapest "hardware upgrade" in ML is noticing that.

Architectural Trade-offs & Production Realities

Architectural Advantages

  • Dynamic batching: throughput without client changes; the delay budget bounds worst-case added latency.
  • Continuous batching: near-full GPU occupancy for ragged-length LLM traffic; 2-4x throughput gains.
  • Queue policies convert overload into defined failures instead of p99 collapse.

Trade-offs & Constraints

  • Every queue-delay microsecond lands on latency at low traffic — batching only pays when arrivals overlap.
  • Padding waste on ragged inputs; needs length bucketing or sequence batching to fix.
  • Scheduler complexity (prefill/decode mixing, preemption) is real operational surface.
Production Implementation in Big Tech
Meta / Orca lineage -> vLLM adopters (e.g., cloud LLM providers)• Serving chat workloads at 100k+ tokens/sec per node

Chat traffic is wildly ragged (8-token answers next to 2k-token coding tasks). Providers run continuous-batching schedulers with chunked prefill: each iteration recomposes the decode set from active + waiting queues, keeping token efficiency above ~90% where per-request static batching idles GPU slots on the longest sibling.

Staff+ Engineering Takeaways

  • Batching amortizes weight loads and kernel launches; throughput scales near-linearly until the GPU saturates, then latency degrades.
  • Dynamic batching lets the server hold arrivals up to a bounded queue delay to form batches — a literal latency/throughput dial.
  • Continuous batching recomposes the running set after every token step, eliminating idle batch slots for ragged LLM workloads (2-4x gains).
  • Monitor mean batch size, preferred-bucket hit rate, and token efficiency; silent underfilling is the most common capacity waste.

Topic Knowledge Check

Exercise 1 of 3 • Test your architectural comprehension.

Exercise 1 of 30 answered
1

In Triton dynamic batching, what does max_queue_delay_microseconds control?

Rate This Architecture ChapterFeedback & Rating

How clear and actionable was this distributed systems breakdown?