Actor-Critic Methods (A2C / A3C)
Actor-critic splits the agent into two networks: an actor (policy) that chooses actions and a critic (value function) that scores them, converting REINFORCE's jittery whole-episode returns into low-variance, step-by-step TD advantages. A3C parallelized the idea across asynchronous worker threads; A2C is its cleaner, synchronous, vectorized successor.
01.The Problem: Waiting for the Final Score Makes Learning Jittery
From the previous topic: REINFORCE nudges action probabilities by their return G_t. It works, and it suffers.
G_tis a Monte-Carlo sum over all the future dice — the environment's, and the policy's own. One lucky jackpot and every action in the episode gets a "more of this!" nudge. The gradient is unbiased but enormously variable: one batch says "go left," the next screams "go right."- Baselines (subtract
V(s)) remove the "this state was just good" inflation — but you still have to wait until the episode ends to knowG_tat all.
Meanwhile, topic 200 showed a calmer teacher: the TD error δ = r + γV(s') − V(s) — one step of reality, bootstrap the rest, update immediately.
What if the policy-gradient engine (topic 203) ran on TD fuel (topic 200) instead of Monte-Carlo fuel?
That combination is actor-critic — and it became the dominant paradigm of modern RL.
Actor–Critic with TD Advantage 🎭
Actor–Critic with TD Advantage 🎭
The ACTOR maps state to an action distribution (policy gradient). The CRITIC estimates V(s) and forms the low-variance TD advantage delta = r + gamma*V(s') - V(s). The actor ascends grad log pi * delta; the critic regresses toward the same targets. Entropy bonus keeps exploration alive.
Unlock Topic #204: Actor-Critic Methods (A2C / A3C)
You are viewing a preview. The full in-depth technical walkthrough, worked derivations, and code notebooks for this concept, along with self-assessment quizzes, are available with Pro or Lifetime Access.
Failure modes, high-throughput bottlenecks, and real FAANG implementation decisions.
Interactive system topology diagrams, live parameter simulators, and downloadable SVG charts.
Staff-level multiple-choice quiz questions with instant feedback and answer explanations.
Firebase Google authentication automatically syncs your completed topics and quiz scores.
How clear and actionable was this distributed systems breakdown?