MLOps & Deployment
Getting models into production and keeping them healthy:
All Topics in Phase 14
0 of 17 completedA deep learning framework is the kitchen where your model gets trained. PyTorch won with "run every line immediately" (eager) ergonomics; TensorFlow bet early on "compile a whole recipe first" (graphs) for production speed — and both now do both: torch.compile and XLA recover graph performance without rewriting code. This topic compares their execution models, ecosystems in 2024-2026, distribution and checkpointing primitives, and why your framework choice matters far less at serving time than at training time.
A GPU is not a fast CPU — it is a ten-thousand-handed slow-cook machine whose speed is set by how far it must walk to the pantry. This topic explains how GPUs actually execute work (SIMT warps, occupancy, divergence), the memory hierarchy from HBM up to registers, why coalescing and kernel fusion are the real performance levers, what streams/MPS/MIG do for sharing a GPU, and where modern abstraction layers like Triton and torch.compile sit above raw CUDA.
A model that scores in a notebook is an ingredient; a model that answers 2,000 requests per second under an SLO is a restaurant. This topic is the anatomy of an inference serving system: the interface (HTTP/gRPC), versioned artifact loading, scheduling and dynamic batching, GPU allocation and multi-model tenancy, warm-up and health checks, autoscaling on concurrency instead of CPU, and the 2024-2026 server landscape (Triton, TF Serving, TorchServe, vLLM) versus DIY FastAPI.
ONNX is the USB-C of machine learning: one self-contained model file — a computation graph plus weights, drawn from a versioned dictionary of standard operators — that PyTorch or TensorFlow exports and ONNX Runtime, TensorRT, CoreML, or a browser then executes. Covers opsets, export workflows, dynamic shapes, quantization, custom ops, and where interchange still breaks down.
TensorRT is a compiler, not just a runtime: a slow build phase specializes a trained model for one exact NVIDIA GPU — fusing layers, auto-tuning kernels, mixing FP16/INT8/FP8 precisions with calibration — and the resulting engine then runs with millisecond-class latency. Covers the build-vs-run split, why engines are hardware-locked, dynamic-shape profiles, plugins, and TensorRT-LLM for autoregressive serving.
A GPU is like a bus: sending one request per trip wastes almost everything. Batching converts waiting into throughput — static (client assembles fixed batches), dynamic (server collects arrivals within a small delay budget), and continuous/in-flight (LLM decode pools where sequences join and leave every token step). This topic covers the economics, the padding-waste and tail-latency math each strategy trades, and the Triton/vLLM knobs that control them.
Latency (what the user feels) and throughput (what the bill reflects) are two different currencies: batching buys throughput cheaply only until the GPU saturates — past that point every extra request/second is extracted directly from latency. This topic does the math: Little's law for fleet sizing, p99 budgets, the TTFT vs inter-token split for LLMs, and cost-per-request as the tie-breaker.
What actually caps LLM serving throughput is not raw math — it is memory bookkeeping. Every token in flight needs its KV cache stored; naive one-big-buffer-per-request allocation wastes 60-80% of GPU memory. PagedAttention borrows the operating system's paging idea (block tables, demand paging, copy-on-write) to make KV memory dense and shareable, which is what lets vLLM run continuous batching safely and stack prefix caching, chunked prefill, and speculative decoding on top.
A spreadsheet of hyperparameters is not an experiment record. A tracked run must capture enough to re-execute and audit it: config + code commit + immutable data version + metric curves + artifacts + lineage + cost. This topic covers what to record, how MLflow and W&B implement it (including the 2025 OpenTelemetry GenAI convergence), and how trackers anchor registry promotion and audit.
Git cannot hold 500 GB of training data — but the same discipline still applies. DVC content-hashes your data and pipelines into tiny pointer files that Git tracks, keeps the bytes in S3/GCS remotes, and lets one Git commit pin the exact dataset+pipeline state a model was trained on. Covers the command workflow, where DVC fits versus warehouses/feature stores, and operational practice.
Feature stores as the consistency layer between training and serving: one feature definition computed for both the offline warehouse (training/backfill) and a low-latency online store (Redis/DynamoDB), with point-in-time correct joins killing training/serving skew.
The model registry as the source of truth between ML and production: versioned model records with stages/aliases, lineage to training runs and data versions, approval workflows, artifact integrity, and how serving consumes it for canaries, rollback, and audit.
A normal CI/CD pipeline tests code. An ML pipeline must also test the data, the model itself, and the whole serving system — through automated gates: data contracts, validation thresholds, artifact parity checks, staged rollout (shadow, canary), and instant rollback. That chain of gates is what turns a promising "candidate" model into a trustworthy "champion" you dare to put in front of users.
A better notebook score does not prove a better product. This topic is the evidence ladder models climb into production: shadow (dark) traffic as the zero-risk first gate, canary and A/B experiments with a sound unit of randomization and guardrail metrics, and the statistical traps — peeking, sample-ratio mismatch, offline-online gaps — that ship fake wins.
A broken server throws errors; a broken model confidently makes worse predictions and says nothing. This topic is the post-deployment loop: the four signal families to watch (system, input, output, outcome), drift detectors you can run without labels (PSI, KS, embedding distance), alert tiers, and the remediation path that feeds retraining through the CI/CD gates.
Federated learning flips the usual pipeline upside down: instead of dragging everyone's data to one server, the model visits the data. Devices or hospitals train locally, send back only their learned updates, and a server averages those updates into a better shared model — raw data never leaves home. This topic covers FedAvg, why non-IID data is the hard part, secure aggregation, differential privacy, communication costs, and when federated learning is (or is not) the right architecture.
Cloud serving sends data to the model; edge serving sends the model to the data — running inference on the phone, camera, or car itself. This topic covers the four honest reasons to do it (latency, cost, privacy, offline), the compression toolbox that shrinks a model 50-500x (pruning, distillation, quantization), the mobile runtimes (LiteRT, Core ML, ExecuTorch, ONNX Runtime Mobile, llama.cpp), and the firmware-grade update and telemetry discipline that keeps it all healthy.