PHASE 9 CURRICULUM

Retrieval, Memory & Agents

Progress0 of 14 (0%)

How LLMs become systems:

Key Architectural Domains & Syllabus
retrieval-augmented generation end to end
chunking strategies
vector databases and ANN indexes
dense vs sparse and hybrid retrieval
reranking with cross-encoders
function calling and tool use
agent loops (ReAct, planning with CoT/ToT/GoT)
short- and long-term memory designs
the Model Context Protocol
agentic feedback loops used in production copilots
14 In-Depth Topics ~112 Minutes Reading Time Interactive Quizzes & Assessments

All Topics in Phase 9

0 of 14 completed

RAG means letting the model look things up instead of guessing: fetch the right passages from your own documents, hand them to the LLM, and make it answer only from that evidence — with citations. It fixes stale knowledge, hallucination, and access control in one architecture.

13 min read•3 Quiz Questions

A retrieval embedding gives every piece of text a set of coordinates in a learned "meaning map", so texts with similar meaning land close together. Search then becomes simple geometry: rank by angle. This topic covers how those maps are trained, measured, and chosen.

12 min read•3 Quiz Questions

A vector database stores embedding coordinates and answers "what is nearest to this query?" fast enough for live traffic. It does so with approximate indexes (HNSW graphs, IVF-PQ clusters) that trade a little accuracy for 100–1000x speed, plus metadata filtering — the feature that decides whether your RAG is usable or a data leak.

13 min read•3 Quiz Questions
#164Semantic SearchAdvancedFREE

Semantic search matches what a query MEANS, not which letters it shares, by comparing query and content as coordinates in one learned meaning-space. In production it is not one model call but a funnel: query understanding, hybrid multi-recall, fusion, reranking, and business signals.

12 min read•3 Quiz Questions

There are two ways to find documents: match exact words (sparse — BM25 over an inverted index) or match learned meaning (dense — DPR dual-encoder vectors). Each fails where the other shines, so production systems run both and merge the two ranked lists with Reciprocal Rank Fusion.

12 min read•3 Quiz Questions

Chunking means slicing documents into retrieval-sized pieces, and the slice you choose decides what your search can ever find: one vector can only carry about 512 tokens of meaning, but every cut can sever a cross-reference. This topic compares fixed, recursive, structural, semantic, late, and parent-child schemes.

12 min read•3 Quiz Questions

Reranking means taking the ~100 candidates a fast first-stage search returns and re-scoring them carefully, pair by pair, with a model that lets the query and document attend to each other. It is the single highest-leverage quality upgrade in a RAG or search stack — paid for by only ever running it on a small pool.

12 min read•3 Quiz Questions

Function calling means the model writes a structured JSON request instead of prose — "call get_order_status with order_id ORD-123" — and YOUR code validates it, runs it as the user, and feeds the result back. The model only proposes; the runtime executes. This topic covers the contract, the reliability levers, and the security boundary.

13 min read•3 Quiz Questions

An agent is an LLM running in a loop that decides its own next action — call a tool, look at the result, decide again — until the task is done or a budget says stop. The hard, valuable part is everything wrapped around that loop: context engineering, budgets, guardrails, and traces. This topic maps the loop, the workflows-vs-agents decision, and the 2025 production skeleton.

13 min read•3 Quiz Questions

ReAct makes a model alternate Thought → Action → Observation: reason one step, do one step, look at real feedback, reason again. That loop kills hallucinated intermediate facts and became the skeleton of every modern AI agent.

13 min read•3 Quiz Questions

How LLMs deliberate: Chain-of-Thought writes one reasoning line, Tree-of-Thoughts explores branches and backtracks, Graph-of-Thoughts merges many partial answers. This topic covers how each works, what each costs, and what reasoning models changed in 2024–2025.

14 min read•3 Quiz Questions

An LLM forgets everything outside its context window. Agent memory means managing that hot window (compaction) and building a cold pipeline outside it: write facts down, retrieve them by recency/importance/relevance, and forget on purpose.

14 min read•3 Quiz Questions

MCP is a JSON-RPC client/server standard that lets any AI app consume tools, data, and prompt templates through one universal plug — like USB-C for models. This topic covers the roles, the three primitives, transports, lifecycle, and the security problems it creates.

14 min read•3 Quiz Questions

One-shot generation commits to its first answer. Feedback loops add the missing check: verify the output (ideally against tests or schemas, not vibes), write down what went wrong, and retry — then turn production failures into permanent regression tests.

14 min read•3 Quiz Questions