Phase 19 Interactive Simulators(5)
Browse all 5 labs →Feature Store & GPU Inference Serving
Tune Triton dynamic batching and the vLLM KV-cache pool to see how GPU utilization, latency and VRAM capacity interact.
Assembled batch
4 rows
over 2ms window @ 2,000 QPS
GPU utilization
100%
Tensor Cores well fed
p50 latency
11.7 ms
queue + GPU + 3ms Feast fetch + 5ms net
H100s needed
2
$16.1 per 1M inferences
LLM KV-Cache pool (LLaMA-3 70B, GQA: 80 layers x 8 KV heads x d=128)
KV per token
0.625 MB
FP16
KV per session
2.68 GB
at 4,096 tokens
VRAM headroom
90 GB
2xH100 pool minus 70 GB weights
Max sessions
33
16-token blocks, on-demand
Batch size 1 leaves >85% of the H100 SMs idle: a 2ms window at 2,000 QPS assembles ~4-row batches and pushes utilization upward, while the offline Point-in-Time store keeps training identical to the online Redis features the gateway joins in <3ms. A single 4,096-token FP16 session consumes ~2.7 GB of VRAM, so legacy contiguous allocation collapses capacity long before PagedAttention does.
Emerging Paradigms & Advanced Topics
Modern system architecture extends beyond classical stateless CRUD services.
All Topics in Phase 19
0 of 5 completedArchitecting high-throughput, low-latency machine learning and LLM serving systems: unified online/offline feature stores (Feast/Hopsworks), dynamic batching with Triton Inference Server, KV-cache memory management (PagedAttention/vLLM), and silent data/concept drift detection.
Building multi-stage recommendation funnels at 100M+ catalog scale: Two-Tower dual encoders, Approximate Nearest Neighbor (ANN) vector indexing (HNSW, IVF-PQ, SCaNN), multi-task ranking with Mixture-of-Experts (MMoE), and sub-50ms business re-ranking.
Architecting sub-10ms global applications: Anycast BGP routing, V8 Isolates vs Containers, WebAssembly (Wasm/WASI) sandboxing, Edge KV, Distributed SQLite (D1/Turso), and distributed state synchronization via CRDTs and Durable Objects.
Architecting resilient multi-cloud and hybrid cloud topologies: overcoming data gravity and egress costs, asynchronous CDC data replication, unified Kubernetes control planes, and sidecarless kernel-bypass networking/observability via eBPF and Cilium.
Architecting immutable financial ledgers, Merkle tree anchoring, WORM storage compliance (SEC Rule 17a-4), crypto-shredding for GDPR Right-to-Erasure, Web3 content-addressed storage (IPFS/Merkle DAGs), and transitioning to NIST Post-Quantum Cryptography (ML-KEM/Kyber, ML-DSA/Dilithium).