Proximal Policy Optimization (PPO)
The default on-policy RL algorithm of the 2020s. PPO clips the probability ratio between the new and the old policy, which caps how far one update can move the policy — so each expensive batch of experience can be safely reused for several gradient steps. It is the engine behind ChatGPT-class RLHF.
01.The Problem: One Big Step Can Wreck the Policy
Reinforcement learning agents learn by doing.
An agent plays its policy. That fancy word just means its behaviour rule: given a situation, which action do I pick, and with what probability.
It collects a batch of experience, then nudges the policy so actions that went well become more likely and actions that went badly become less likely.
Here is the catch.
Why is collecting that experience so precious?
Because it is expensive. A robot wears out hardware. A game agent burns simulated hours. An LLM in RLHF has to generate thousands of full answers just so a scorer can rate them.
And that batch was generated by the old policy, π_old.
So ask
What happens if the update step is too large?
Two failures, back to back:
- The policy can collapse into a bad distribution — say, it suddenly only produces one kind of sentence. Exploration dies, and every future batch of data is worse.
- The data you just paid for described the old policy. After a giant step, its gradients are misleading advice about a policy that no longer exists.
This is the classic fragility of vanilla policy gradients and plain actor-critic: one too-large step throws away the batch you just paid to collect.
TRPO (Schulman et al., 2015) fixed this with a hard KL trust-region constraint and a second-order solve — principled, but complex, memory-hungry, and awkward with minibatches.
PPO (Schulman et al., 2017) delivers most of TRPO's safety with a first-order, minibatch-friendly, easy-to-implement objective. It became the de facto default RL algorithm.
PPO-Clip: The Clipped Ratio 🛡️
PPO-Clip: The Clipped Ratio 🛡️
PPO measures how far the new policy has moved from the rollout policy with the probability ratio r = pi_theta/pi_old. It maximizes min(r*A, clip(r,1-eps,1+eps)*A). The clip shuts off improvement gradients when r leaves the trust region, so a single batch can be optimized for multiple epochs without the policy blowing up.
Unlock Topic #205: Proximal Policy Optimization (PPO)
You are viewing a preview. The full in-depth technical walkthrough, worked derivations, and code notebooks for this concept, along with self-assessment quizzes, are available with Pro or Lifetime Access.
Failure modes, high-throughput bottlenecks, and real FAANG implementation decisions.
Interactive system topology diagrams, live parameter simulators, and downloadable SVG charts.
Staff-level multiple-choice quiz questions with instant feedback and answer explanations.
Firebase Google authentication automatically syncs your completed topics and quiz scores.
How clear and actionable was this distributed systems breakdown?