Policy Gradient Methods
Instead of learning a value table and extracting a policy from it, policy gradient tunes the policy itself: roll out the stochastic policy pi_theta, measure returns, and nudge the parameters so actions that earned applause become more probable (theta += alpha * grad J). REINFORCE, the score-function trick, and the baseline-to-advantage bridge are the core math; actor-critic and PPO are next door.
01.The Problem: Value Tables Have a Ceiling
The value-based line (Bellman → TD → Q-learning → DQN) has been a great run. But try to use it for:
- A robotic arm. Its action is a continuous torque vector — there is no list of actions to take an argmax over.
max_aover infinitely many actions is its own unsolved optimization problem. - Poker, or a security guard scheduling patrols. The optimal behavior is deliberately random — bluffing frequencies, patrol mixes. A deterministic greedy policy (what argmax hands you) is exploitable (see rock-paper-scissors, topic 197).
Both want the same thing value methods cannot naturally give: a smooth, stochastic policy you can differentiate.
So flip the strategy:
If the policy is a differentiable function
π_θ, why not just… measure how good its play was, and pushθin the direction that makes good play more likely?
That is the policy gradient. Skip the value table entirely (or keep it only as a helper), parameterize the behavior, and do calculus on it directly.
The Policy Gradient Loop 📈
The Policy Gradient Loop 📈
Unlike value methods, we optimize the policy parameters theta directly: roll out pi_theta, measure reward, and take gradient ASCENT on the expected return. The log-derivative trick lets us compute grad E[reward] without knowing the (unknown) distribution gradient, and a baseline b(s) cuts variance for free.
Unlock Topic #203: Policy Gradient Methods
You are viewing a preview. The full in-depth technical walkthrough, worked derivations, and code notebooks for this concept, along with self-assessment quizzes, are available with Pro or Lifetime Access.
Failure modes, high-throughput bottlenecks, and real FAANG implementation decisions.
Interactive system topology diagrams, live parameter simulators, and downloadable SVG charts.
Staff-level multiple-choice quiz questions with instant feedback and answer explanations.
Firebase Google authentication automatically syncs your completed topics and quiz scores.
How clear and actionable was this distributed systems breakdown?