LLM and Prompt Caching
Every LLM call re-reads its prompt from scratch — unless you cache it. Prompt caching reuses the model's "reading notes" (the KV cache) for the unchanged beginning of your prompt, cutting cost up to 90% and time-to-first-token by ~80% on long prompts. This topic covers what is cached, provider pricing, the prompt layout that actually hits, and the routing behind it.
01.The Problem: Re-reading the Manual Every Morning
Your support assistant has a 30,000-token prompt: system instructions, tool schemas, the product manual, and 40 customer Q&A examples. It runs 5,000 times a day.
Every single call, the model re-reads the entire 30,000 tokens from scratch — doing the full internal math to "understand" the same words it understood a second ago — before answering a one-line question.
So the question becomes
If the first 30,000 tokens are identical every time, why pay to recompute them?
With a normal program the fix is obvious: memoize. LLMs have a version of that too. The model's "reading notes" for a prompt are called the KV cache, and prompt caching means keeping those notes for a prefix you have seen before, so the next call skips straight to the new part.
Done right: up to 90% off the input cost of the cached part and roughly 80% lower time-to-first-token on 100K prompts. Done wrong — one timestamp at the top — a 0% hit rate, and teams wonder why "we enabled caching" and saved nothing.
Carry this analogy through: a book and a set of highlighters. You annotated pages 1-300 of a manual last week. Today someone asked you to start reading from page 1 again, but your highlighter notes are still in the book — you skip to page 301 instantly. The one rule that matters:
Your notes stay valid ONLY if nobody edited an earlier page. Change page 3, and every note from page 3 onward is garbage.
Prefix Caching Is a Layout Problem 🧊
Prefix Caching Is a Layout Problem 🧊
A cache hit requires byte-identical token prefixes routed to memory that still holds their KV. Anything that changes early in the prompt - even a timestamp - invalidates everything after it.
Unlock Topic #264: LLM and Prompt Caching
You are viewing a preview. The full in-depth technical walkthrough, worked derivations, and code notebooks for this concept, along with self-assessment quizzes, are available with Pro or Lifetime Access.
Failure modes, high-throughput bottlenecks, and real FAANG implementation decisions.
Interactive system topology diagrams, live parameter simulators, and downloadable SVG charts.
Staff-level multiple-choice quiz questions with instant feedback and answer explanations.
Firebase Google authentication automatically syncs your completed topics and quiz scores.
How clear and actionable was this distributed systems breakdown?