TOPIC #227Advanced 15 min read

vLLM & PagedAttention: Operating Systems Ideas for LLM Serving

AI
AI & ML Editorial
Report an issue
Key takeawayCore Concept Summary

What actually caps LLM serving throughput is not raw math — it is memory bookkeeping. Every token in flight needs its KV cache stored; naive one-big-buffer-per-request allocation wastes 60-80% of GPU memory. PagedAttention borrows the operating system's paging idea (block tables, demand paging, copy-on-write) to make KV memory dense and shareable, which is what lets vLLM run continuous batching safely and stack prefix caching, chunked prefill, and speculative decoding on top.

01.The Problem: The GPU Is Full, and It Is Not the Math

You run an 8-billion-parameter chat model on an 80 GB GPU. The weights take ~16 GB in fp16. How many people can chat at once?

The beginner guess: "as many as the math allows." The real answer involves a stranger: decode — generating tokens one at a time — is memory-bound, and the memory holding the model's own memory is the bottleneck.

Here is the mechanic. To generate token #501, the model "looks back" at all 500 previous tokens — specifically at each one's key and value vectors. Together they are the KV cache, and every token in every live request needs its KV stored somewhere, then read every step.

So the serving question becomes

Insight

How many tokens' worth of KV can we fit in the GPU's leftover memory — with the least waste?

Pre-vLLM engines answered it terribly: they allocated one contiguous buffer per request, sized for the maximum possible length. The 2023 paper measured the fallout: only 20-38% of KV memory held useful tokens; the rest was reservation waste and external fragmentation. Worse, requests were rejected or blocked even when aggregate free memory sufficed — no single contiguous chunk fitted.

The historical punchline: operating systems had this exact disease in the 1960s, and cured it with a trick called paging.

From Contiguous Reserves to Paged KV Blocks

PRO Architecture Blueprint

From Contiguous Reserves to Paged KV Blocks

Pre-vLLM engines reserved one contiguous KV buffer per request at max length, wasting 60-80% of GPU memory to fragmentation and early reservation. PagedAttention maps each request's logical KV positions to small physical blocks via a block table, enabling reuse and sharing.

From Contiguous Reserves to Paged KV Blocks
100%
Touchpad: Pinch to zoom • Drag to pan
Rendering visual architecture flowchart...
PRO & LIFETIME CURRICULUM

Unlock Topic #227: vLLM & PagedAttention: Operating Systems Ideas for LLM Serving

You are viewing a preview. The full in-depth technical walkthrough, worked derivations, and code notebooks for this concept, along with self-assessment quizzes, are available with Pro or Lifetime Access.

Production Deep Dive

Failure modes, high-throughput bottlenecks, and real FAANG implementation decisions.

Interactive Blueprints

Interactive system topology diagrams, live parameter simulators, and downloadable SVG charts.

Knowledge Assessment

Staff-level multiple-choice quiz questions with instant feedback and answer explanations.

Cross-Device Progress Sync

Firebase Google authentication automatically syncs your completed topics and quiz scores.

Rate This Architecture ChapterFeedback & Rating

How clear and actionable was this distributed systems breakdown?