Memory in Agents: Context Windows, Vector Stores & Episodic Recall
An LLM forgets everything outside its context window. Agent memory means managing that hot window (compaction) and building a cold pipeline outside it: write facts down, retrieve them by recency/importance/relevance, and forget on purpose.
01.The Problem: An Assistant With No Memory At All
Imagine an assistant you have used for a year. It knows you hate aisle seats, your project uses a staging server called VPN-2, and last month you spent three days debugging a weird date bug.
Now open a new chat. It knows nothing. Not your name, not your project, not the bug.
That is not a bug — it is the raw nature of an LLM: the model is stateless. Outside the text currently in front of it (the context window: everything the model gets to see on one run — your question plus everything before it in the conversation), nothing exists. When the session ends, the window is wiped.
Two naive fixes both fail:
- "Just make the window bigger." Windows grew from 128k to over 1M tokens by 2025. But attention quality degrades with length ("lost in the middle" — models recall the start and end of long contexts far better than the middle), every call costs tokens proportional to the window you send, and — decisively — writes still don't happen: pasting more text into a prompt never persists a fact into next week's session.
- "Just retrain the model." Weight updates per user are absurdly expensive and can't be undone.
So the question becomes
How do you give a stateless text-eating machine the feeling of remembering — across turns, across sessions, across months?
The answer has two halves: manage the hot window, and build a cold pipeline outside it.
Agent Memory Architecture
Agent Memory Architecture
One hot working set (the window) fed by and writing back to three cold stores, with an explicit write/recall/forget pipeline.
Unlock Topic #172: Memory in Agents: Context Windows, Vector Stores & Episodic Recall
You are viewing a preview. The full in-depth technical walkthrough, worked derivations, and code notebooks for this concept, along with self-assessment quizzes, are available with Pro or Lifetime Access.
Failure modes, high-throughput bottlenecks, and real FAANG implementation decisions.
Interactive system topology diagrams, live parameter simulators, and downloadable SVG charts.
Staff-level multiple-choice quiz questions with instant feedback and answer explanations.
Firebase Google authentication automatically syncs your completed topics and quiz scores.
How clear and actionable was this distributed systems breakdown?