Reinforcement Learning
Learning by interaction:
All Topics in Phase 12
0 of 13 completedAn MDP is the rule sheet for any sequential decision problem: states, actions, the world's dice roll (transition probabilities), rewards, and a discount factor. Its one big assumption — the Markov property, "the future depends only on the present" — is what makes reinforcement learning math tractable.
A policy pi is the agent's cheat sheet: look at the state, name the action (or the dice-roll over actions). It — not the value function — is the thing RL actually deploys. This topic covers deterministic vs stochastic forms, the Policy Improvement Theorem that makes rewrite-after-rewrite converge, and on-policy vs off-policy learning.
V(s) is the fair price of standing in a state — the long-run reward you can expect from there. Q(s,a) is the price of standing there AND making a specific move. That second number is decision-ready: with Q you can act without any model of the world, and subtracting V from Q gives the advantage that powers modern policy methods.
The Bellman equation turns the scary infinite sum of future rewards into one simple recursion: value of here = immediate reward + discounted value of where I land next. Expectation form scores a fixed policy; optimality form (with a max) yields the best policy — and iterating either one is how dynamic programming, TD learning, and every deep RL backup actually work.
TD learning predicts, observes one step of reality, and corrects — updating today's guess toward the reward you just saw plus tomorrow's guess. It fuses Monte-Carlo honesty with dynamic-programming bootstrapping, giving model-free, online updates that power TD(0), SARSA, eligibility traces, TD(lambda), and the critics of PPO.
Q-learning is the canonical model-free, off-policy way to learn the best action-value: after every step, nudge Q(s,a) toward r + gamma*max Q of the next state, while actually acting epsilon-greedy. The max-in-the-target trick lets a clumsy, exploratory player still learn the optimal policy — provably, and with no model of the world.
DQN is Q-learning scaled to raw pixels: a deep network approximates Q(s,a) over huge state spaces, and two inventions — an experience replay buffer (shuffled flashcards) and a frozen target network (a slowly-updated answer key) — stop the naive combination from diverging. Double DQN, Dueling, Prioritized Replay, and Rainbow are its landmark refinements.
Instead of learning a value table and extracting a policy from it, policy gradient tunes the policy itself: roll out the stochastic policy pi_theta, measure returns, and nudge the parameters so actions that earned applause become more probable (theta += alpha * grad J). REINFORCE, the score-function trick, and the baseline-to-advantage bridge are the core math; actor-critic and PPO are next door.
Actor-critic splits the agent into two networks: an actor (policy) that chooses actions and a critic (value function) that scores them, converting REINFORCE's jittery whole-episode returns into low-variance, step-by-step TD advantages. A3C parallelized the idea across asynchronous worker threads; A2C is its cleaner, synchronous, vectorized successor.
The default on-policy RL algorithm of the 2020s. PPO clips the probability ratio between the new and the old policy, which caps how far one update can move the policy — so each expensive batch of experience can be safely reused for several gradient steps. It is the engine behind ChatGPT-class RLHF.
When the optimized reward diverges from the intended objective, agents exploit the gap. Reward hacking, specification gaming, and Goodhart's law are the central safety failure modes of RL and RLHF in 2024-2026.
The fundamental trade-off: exploit the known-best action or explore to reduce uncertainty. Epsilon-greedy, upper-confidence bound, Boltzmann softening, Thompson sampling, and intrinsic-reward/RLHF-era exploration strategies.