ReAct: Reasoning and Acting Interleaved
ReAct makes a model alternate Thought → Action → Observation: reason one step, do one step, look at real feedback, reason again. That loop kills hallucinated intermediate facts and became the skeleton of every modern AI agent.
01.The Problem: A Brain That Cannot Look Anything Up
Ask an AI assistant:
"Who founded the company that acquired Instagram?"
Easy? Maybe. The model has to chain two facts: who acquired Instagram → who founded that company. Each hop is a chance to be wrong — and real user questions chain three, four, five hops over APIs, databases, and the web.
There are two naive ways to handle this, and each fails in its own way.
Way 1: think, never touch. Chain-of-thought prompting (topic 171 — in one plain sentence: making the model write out its reasoning step by step before answering) plans beautifully. But it never checks the outside world. Every fact must come out of the model's frozen training memory. When a fact is missing, the model does not shrug — it invents one, fluently and confidently. That is hallucination, and it is CoT's worst failure on multi-hop questions: the intermediate steps are invented, so the final answer is built on sand.
Way 2: touch, never think. At the other extreme, a pure action policy just does things: search, click, call the API. Every step is grounded in reality — but there is no plan. Actions become greedy: whatever query looks reasonable right now, chosen hop by hop. Get the first query wrong and everything after it is wasted.
So the question becomes
Can the model plan out loud AND ground each step in real feedback — inside the same loop?
ReAct (Yao et al., ICLR 2023, arXiv:2210.03629) answered yes, with one deceptively simple prompt pattern.
The ReAct Trajectory
The ReAct Trajectory
Reasoning traces condition future actions; observations correct the reasoning. Neither chain-of-thought nor act-only alone stays accurate on multi-hop QA.
Unlock Topic #170: ReAct: Reasoning and Acting Interleaved
You are viewing a preview. The full in-depth technical walkthrough, worked derivations, and code notebooks for this concept, along with self-assessment quizzes, are available with Pro or Lifetime Access.
Failure modes, high-throughput bottlenecks, and real FAANG implementation decisions.
Interactive system topology diagrams, live parameter simulators, and downloadable SVG charts.
Staff-level multiple-choice quiz questions with instant feedback and answer explanations.
Firebase Google authentication automatically syncs your completed topics and quiz scores.
How clear and actionable was this distributed systems breakdown?