TOPIC #226Intermediate 13 min read

Latency vs Throughput: The Inference SLO Math

AI
AI & ML Editorial
Report an issue
Key takeawayCore Concept Summary

Latency (what the user feels) and throughput (what the bill reflects) are two different currencies: batching buys throughput cheaply only until the GPU saturates — past that point every extra request/second is extracted directly from latency. This topic does the math: Little's law for fleet sizing, p99 budgets, the TTFT vs inter-token split for LLMs, and cost-per-request as the tie-breaker.

01.The Problem: Two Different Currencies

A user cares about how long their request takes. Your CFO cares about how many requests each GPU serves per second. These are two different currencies:

  • Latency is per-request time (what a user feels).
  • Throughput is requests (or tokens) per second (what the bill reflects).

Beginner intuition hopes "just make the server faster and both improve." The reality is a three-act trade:

  • Act 1 — low load, batch=1: minimum latency and minimum throughput. The GPU sits mostly idle; every step is dominated by weight reads and kernel launches.
  • Act 2 — grow the batch: throughput rises near-linearly while the GPU is not saturated and latency stays roughly flat — one fused step costs little more than a single pass.
  • Act 3 — past saturation: additional throughput comes directly out of latency. Queueing delay grows, often superlinearly as utilization ρ approaches 1 (the M/M/1 queueing intuition: wait ~ ρ/(1-ρ) — at 90% utilization the wait is nine times the service time).

So the design question is not "how do I get both?" It is

Insight

Where on this curve do we choose to stand — and how do we prove that spot in numbers?

PRO & LIFETIME CURRICULUM

Unlock Topic #226: Latency vs Throughput: The Inference SLO Math

You are viewing a preview. The full in-depth technical walkthrough, worked derivations, and code notebooks for this concept, along with self-assessment quizzes, are available with Pro or Lifetime Access.

Production Deep Dive

Failure modes, high-throughput bottlenecks, and real FAANG implementation decisions.

Interactive Blueprints

Interactive system topology diagrams, live parameter simulators, and downloadable SVG charts.

Knowledge Assessment

Staff-level multiple-choice quiz questions with instant feedback and answer explanations.

Cross-Device Progress Sync

Firebase Google authentication automatically syncs your completed topics and quiz scores.

Rate This Architecture ChapterFeedback & Rating

How clear and actionable was this distributed systems breakdown?