PHASE 12 CURRICULUM

Reinforcement Learning

Progress0 of 13 (0%)

Learning by interaction:

Key Architectural Domains & Syllabus
Markov decision processes
reward design
dynamic programming
Monte Carlo and temporal-difference learning
Q-learning and DQN
policy gradients and actor-critic methods
PPO and modern policy optimization
offline RL
the role of RL in aligning and reasoning with LLMs (RLHF, DPO, GRPO, and RL with verifiable rewards as used by o-series and DeepSeek-R1 style models)
13 In-Depth Topics ~104 Minutes Reading Time Interactive Quizzes & Assessments

All Topics in Phase 12

0 of 13 completed

An MDP is the rule sheet for any sequential decision problem: states, actions, the world's dice roll (transition probabilities), rewards, and a discount factor. Its one big assumption — the Markov property, "the future depends only on the present" — is what makes reinforcement learning math tractable.

12 min read•3 Quiz Questions

A policy pi is the agent's cheat sheet: look at the state, name the action (or the dice-roll over actions). It — not the value function — is the thing RL actually deploys. This topic covers deterministic vs stochastic forms, the Policy Improvement Theorem that makes rewrite-after-rewrite converge, and on-policy vs off-policy learning.

11 min read•3 Quiz Questions

V(s) is the fair price of standing in a state — the long-run reward you can expect from there. Q(s,a) is the price of standing there AND making a specific move. That second number is decision-ready: with Q you can act without any model of the world, and subtracting V from Q gives the advantage that powers modern policy methods.

11 min read•3 Quiz Questions
#199The Bellman EquationAdvancedFREE

The Bellman equation turns the scary infinite sum of future rewards into one simple recursion: value of here = immediate reward + discounted value of where I land next. Expectation form scores a fixed policy; optimality form (with a max) yields the best policy — and iterating either one is how dynamic programming, TD learning, and every deep RL backup actually work.

12 min read•3 Quiz Questions

TD learning predicts, observes one step of reality, and corrects — updating today's guess toward the reward you just saw plus tomorrow's guess. It fuses Monte-Carlo honesty with dynamic-programming bootstrapping, giving model-free, online updates that power TD(0), SARSA, eligibility traces, TD(lambda), and the critics of PPO.

12 min read•3 Quiz Questions
#201Q-LearningAdvancedFREE

Q-learning is the canonical model-free, off-policy way to learn the best action-value: after every step, nudge Q(s,a) toward r + gamma*max Q of the next state, while actually acting epsilon-greedy. The max-in-the-target trick lets a clumsy, exploratory player still learn the optimal policy — provably, and with no model of the world.

12 min read•3 Quiz Questions
#202Deep Q-Network (DQN)Advanced PRO

DQN is Q-learning scaled to raw pixels: a deep network approximates Q(s,a) over huge state spaces, and two inventions — an experience replay buffer (shuffled flashcards) and a frozen target network (a slowly-updated answer key) — stop the naive combination from diverging. Double DQN, Dueling, Prioritized Replay, and Rainbow are its landmark refinements.

12 min read•3 Quiz Questions
#203Policy Gradient MethodsAdvanced PRO

Instead of learning a value table and extracting a policy from it, policy gradient tunes the policy itself: roll out the stochastic policy pi_theta, measure returns, and nudge the parameters so actions that earned applause become more probable (theta += alpha * grad J). REINFORCE, the score-function trick, and the baseline-to-advantage bridge are the core math; actor-critic and PPO are next door.

12 min read•3 Quiz Questions

Actor-critic splits the agent into two networks: an actor (policy) that chooses actions and a critic (value function) that scores them, converting REINFORCE's jittery whole-episode returns into low-variance, step-by-step TD advantages. A3C parallelized the idea across asynchronous worker threads; A2C is its cleaner, synchronous, vectorized successor.

12 min read•3 Quiz Questions

The default on-policy RL algorithm of the 2020s. PPO clips the probability ratio between the new and the old policy, which caps how far one update can move the policy — so each expensive batch of experience can be safely reused for several gradient steps. It is the engine behind ChatGPT-class RLHF.

13 min read•3 Quiz Questions

When the optimized reward diverges from the intended objective, agents exploit the gap. Reward hacking, specification gaming, and Goodhart's law are the central safety failure modes of RL and RLHF in 2024-2026.

12 min read•3 Quiz Questions

The fundamental trade-off: exploit the known-best action or explore to reduce uncertainty. Epsilon-greedy, upper-confidence bound, Boltzmann softening, Thompson sampling, and intrinsic-reward/RLHF-era exploration strategies.

13 min read•3 Quiz Questions

Many learning agents in one shared environment. Non-stationarity, credit assignment, and the CTDE principle; cooperative, competitive and mixed games; QMIX, MAPPO, and self-play that produced superhuman teaming.

13 min read•3 Quiz Questions