ML Feature Store & Model Serving Lab (Interactive)
Tune QPS, batch window, and KV-cache precision to watch GPU utilization, latency and VRAM capacity react. Models Triton dynamic batching and vLLM PagedAttention: batch assembly math sets utilization and p50 latency, while context length and quantization set KV-cache VRAM.
Feature Store & GPU Inference Serving
Tune Triton dynamic batching and the vLLM KV-cache pool to see how GPU utilization, latency and VRAM capacity interact.
Assembled batch
4 rows
over 2ms window @ 2,000 QPS
GPU utilization
100%
Tensor Cores well fed
p50 latency
11.7 ms
queue + GPU + 3ms Feast fetch + 5ms net
H100s needed
2
$16.1 per 1M inferences
LLM KV-Cache pool (LLaMA-3 70B, GQA: 80 layers x 8 KV heads x d=128)
KV per token
0.625 MB
FP16
KV per session
2.68 GB
at 4,096 tokens
VRAM headroom
90 GB
2xH100 pool minus 70 GB weights
Max sessions
33
16-token blocks, on-demand
Batch size 1 leaves >85% of the H100 SMs idle: a 2ms window at 2,000 QPS assembles ~4-row batches and pushes utilization upward, while the offline Point-in-Time store keeps training identical to the online Redis features the gateway joins in <3ms. A single 4,096-token FP16 session consumes ~2.7 GB of VRAM, so legacy contiguous allocation collapses capacity long before PagedAttention does.
How It Works Under the Hood
Serving ML and LLM inference is a packing problem. A batch of 1 leaves over 85% of an H100's streaming multiprocessors idle, so Triton-style gateways wait a 1-2ms window to merge queued tensors into one batch, trading deterministic queue delay for 8-15x throughput. For LLMs the binding constraint is memory: a 4,096-token FP16 KV-cache burns ~2.7 GB per session, and contiguous worst-case allocation wastes 60-80% of VRAM to fragmentation, which PagedAttention blocks and quantization recover.
Core Architectural Principles
- Batch size assembles from arrivals inside the max_queue_delay window, capped by max_batch_size.
- Per-token KV-cache bytes = 2 x 2 x layers x KV heads x d_head x precision, so quantization halves VRAM.
- Feature fetch from the online Redis store joins each request in under 3ms with point-in-time-correct training parity.
Open by allocating the end-to-end latency budget: 3ms feature fetch, batching window plus GPU forward pass, 5ms network. Then show hardware literacy: dynamic batching with a 2ms timeout lifts GPU utilization from ~12% past 90%, and for LLMs compute KV-cache bytes per session to size concurrency, citing PagedAttention as the fragmentation fix. Mention PSI-based drift alerts to close.
Longer batching windows and larger batches maximize GPU utilization and cut cost per inference, but add queue latency and risk KV-cache OOM at high LLM concurrency.