Policy: The Agent's Decision Rule
A policy pi is the agent's cheat sheet: look at the state, name the action (or the dice-roll over actions). It — not the value function — is the thing RL actually deploys. This topic covers deterministic vs stochastic forms, the Policy Improvement Theorem that makes rewrite-after-rewrite converge, and on-policy vs off-policy learning.
Policy in the Agent–Environment Loop 🧭
The policy pi is the agent's behavior: it reads the state and emits (or samples) an action. The environment then returns the next state and reward, closing the loop the policy is trained inside.
01.The Problem: You Can Score Positions — You Still Have to Move
The previous topic (MDP) gave us the rule sheet of a board game: states, actions, the world's dice roll, points, and a discount factor.
Now you are the player.
Given this board, what do I do?
That question needs an answer for every situation the game can put you in — every state, every turn, for the whole episode. The answer is called the policy, and it is the object reinforcement learning is ultimately trying to produce.
Two wrinkles trip up beginners right away:
- The world is random (the die behind the board). So "what should I do" is often really "what should I do, with what probabilities."
- A single move's outcome tells you almost nothing — a good decision can still lose to bad luck. So the agent usually learns long-run scores first (next topic: value functions) and derives the rule from them.
This topic is about the rule itself: what it is, its two flavors, the theorem proving that rewriting it makes it better, and the two schools of how it gets learned.
02.The Idea in Plain Words: A Cheat Sheet with One Line per Situation
A policy π is the agent's decision-making rule — a mapping from states to actions.
The deterministic version, one line written plainly:
π: S → A, a = π(s)
"In state s, do action a. Full stop."
The stochastic version maps each state to a probability distribution over actions:
π(a | s) = Pr(A_t = a | S_t = s)
"In state s, roll the dice: this action 70% of the time, that one 30%."
Read π(a|s) out loud as "the probability the policy plays action a when it sees state s."
Carried-through analogy: the policy is a cheat sheet with one line per situation. You land on a board square (state), you consult the line for that square, and it names your move. A deterministic policy writes exactly one move per line. A stochastic policy writes "flip a coin: heads go here, tails go there."
And here is the key idea that trips up beginners: the policy, not the value function, is what the agent ultimately deploys. Value functions (topic 198) are an intermediate computation used to find a good cheat sheet. When your robot ships to a customer, what runs on the factory floor is π.
03.Deterministic vs Stochastic: One Named Move, or Dice
Both flavors are legal. Each one wins somewhere.
- Deterministic policies pick exactly one action per state. They are optimal in fully observable, discrete MDPs, and cheaper to store and execute — a trained DQN acts greedily, so deployment is just
argmax_a Q(s,a)per state. - Stochastic policies output a probability distribution over actions. They are not optional, for two reasons:
- Math: gradient-based policy methods (topic 203) need a smooth probability distribution to differentiate through. A hard argmax has no slope to climb.
- Strategy: in partially observable or adversarial games, any fixed rule is exploitable. Rock-paper-scissors has no good deterministic policy — see the worked example below.
How neural networks actually represent policies in deep RL:
- Softmax / Categorical over discrete actions: the network outputs one score (logit) per action, and
π_θ(a|s) = softmax(f_θ(s))_aturns those scores into probabilities. - Gaussian for continuous control: the network predicts a mean
μ_θ(s)and a log-variance; the action is sampled fromN(μ_θ(s), σ²). This is how robot torques and drone tilts get expressed. - Deterministic policy gradient (DDPG/TD3): the network outputs a single continuous action
μ_θ(s)directly — no sampling at deployment time.
04.A Worked Example: Rock-Paper-Scissors and One Softmax Line
Play rock-paper-scissors against an opponent who studies your habits.
A deterministic policy "always play Rock" is perfectly predictable: the opponent plays Paper forever and you lose forever. The cheat sheet line is the leak. In fact every deterministic line is exploitable in this game.
Now a stochastic policy: uniform dice on every turn.
π(Rock) = π(Paper) = π(Scissors) = 1/3
Expected payoff versus any opponent: exactly 0. You can never be taken advantage of — the distribution shields you, not the order of moves.
Same idea, one softmax with numbers. Suppose the policy network looks at a state and outputs logits (2, 0) for two actions. Softmax exponentiates and normalizes:
e² ≈ 7.39ande⁰ = 1, so the probabilities are(7.39 / 8.39, 1 / 8.39) ≈ (0.88, 0.12).- Push the logits further apart —
(20, 0)— and the distribution becomes ≈(1.0, 0.0): greedy-looking behavior, while still a smooth, differentiable object during training. - "Temperature" is exactly the dial that flattens or sharpens one cheat-sheet line. Low temperature = nearly deterministic; high = dice.
05.Visual Intuition: The Cheat Sheet Inside the Loop
Every trained agent is just this cycle spinning:
codestate s_t │ look up the line for s_t ▼ ┌────────────────────────────────────┐ │ POLICY π │ │ deterministic: π(s) → one action │ │ stochastic: π(a|s) → dice roll │ └────────────────────────────────────┘ │ play a_t ▼ ENVIRONMENT (P rolls the hidden die, R pays points) │ returns s_{t+1}, r_{t+1} └────────────► back to the top, forever (or until terminal)
The mermaid diagram above draws the same loop: the policy reads states and emits (or samples) actions; the environment returns the next state and reward. Training happens inside this loop; deployment is only this loop, with the cheat sheet frozen.
Short version:
- The policy is the behavior.
- The value function is the scorecard used to rewrite the behavior.
- You ship the behavior.
06.The Policy Improvement Theorem: Rewriting Never Hurts
Why does RL converge to anything good at all? The Policy Improvement Theorem is the guarantee.
If my cheat sheet is imperfect, can rewriting one line ever make things worse?
Not if you rewrite each line to the best action according to the scorecard of your current sheet. Formally: given policy π and its action-values Q_π(s,a), build a new policy π' that acts greedily:
π'(s) = argmax_a Q_π(s,a)
Then π' is at least as good as π in every state:
v_π'(s) ≥ v_π(s) for all s
Intuition: Q_π(s,a) already means "take action a now, then keep following π afterward." If you replace the dice-roll over actions with the highest-scoring action, you cannot lose — you removed exactly the actions that were dragging the average down.
Now alternate two steps forever:
- Policy evaluation — score the current cheat sheet (compute
v_πandQ_π). - Policy improvement — rewrite every line greedily against those scores.
For a finite MDP this provably converges to the optimal policy π*, because value can only rise and there are finitely many policies. The loop is called policy iteration, and it is the ancestor of almost every RL algorithm.
The code below is step 2 in its purest form: one argmax per row of a learned Q-table. Read row by row — state 0 scores [2.1, 0.4, 1.9] across three actions, so the greedy line says "action 0"; state 1's best is action 1 (3.3); state 2's best is action 0 (5.0). The whole deployable cheat sheet falls out as [0, 1, 0].
import numpy as np
# Q[s, a] learned by some method (e.g. Q-learning, topic 201)
Q = np.array([[2.1, 0.4, 1.9], # state 0
[0.7, 3.3, 1.1], # state 1
[5.0, 2.2, 0.5]]) # state 2
def greedy_policy(Q):
"""Deterministic policy improvement: pi(s) = argmax_a Q(s, a)."""
return np.argmax(Q, axis=1)
print("greedy pi:", greedy_policy(Q)) # -> [0, 1, 0]07.On-Policy vs Off-Policy: Who Gathered the Data, Who Is It About
How a policy is learned splits the field in two. The one question to ask:
Is the data generated by the same policy the data is about?
- On-policy: the agent learns only from data generated by the current policy itself. If
πchanges, all older samples are stale and get thrown away. Examples: SARSA, A2C, PPO. Safe and simple, but sample-hungry. - Off-policy: the agent learns about a target policy
πwhile behaving with a different behavior policyμ— usually ε-greedy (mostly exploit, sometimes random). Examples: Q-learning, DQN. Experience is reusable, at the cost of higher variance and a divergence risk when neural nets join in.
codebehavior policy μ ──plays──► data (s, a, r, s') │ ▼ is used to learn about target policy π (may be different!) μ == π → on-policy μ ≠ π → off-policy
Why does the distinction matter? Because logs, human demonstrations, old checkpoints, and other agents' play are all someone else's policy. Off-policy learning is what lets an agent mine them without replaying their mistakes live.
08.Why AI Cares: The LLM Is the Policy
Modern alignment runs make the abstraction concrete: in RLHF, a ChatGPT-class language model is a stochastic policy π_θ(a|s).
- The state is the prompt plus the tokens generated so far.
- The action is the next token, sampled from the model's softmax distribution over the vocabulary — exactly the categorical parameterization from Section 3.
- PPO (topic 205) updates
θto maximize a reward-model score while a KL penalty keeps the cheat sheet anchored near the reference model.
So when someone says "sampling with temperature," "greedy decoding," or "the model became more deterministic after fine-tuning," they are talking about one line of this policy's cheat sheet getting sharper or flatter. The value function you often see in the same pipeline? That is the scorecard — topic 198 — not the thing that ships.
Architectural Trade-offs & Production Realities
Architectural Advantages
- Stochastic policies are differentiable, enabling direct gradient optimization.
- Deterministic greedy policies maximize value and are cheap to store and run at deployment.
- The improvement theorem gives a convergence guarantee for policy iteration.
Trade-offs & Constraints
- Deterministic policies in partially observable or adversarial games are exploitable.
- On-policy learning throws away samples whenever the policy updates.
- Policy search has high gradient variance without baselines and advantages.
In RLHF the LLM is a stochastic policy pi_theta(a|s): the state is the prompt + tokens so far, the action is the next token sampled from the model's softmax distribution. PPO updates this policy to maximize a reward-model score while a KL penalty keeps it near the reference.
Staff+ Engineering Takeaways
- A policy pi maps states to actions (deterministic) or to action probabilities (stochastic) — a cheat sheet with one line per situation.
- The policy is the deployable artifact; value functions are an intermediate step for learning it.
- Stochastic policies are required for differentiable policy-gradient optimization, and for unexploitable play in adversarial games like rock-paper-scissors.
- The Policy Improvement Theorem guarantees that acting greedily w.r.t. Q_pi is never worse than pi, so evaluate-and-rewrite loops (policy iteration) converge to pi*.
- On-policy learns only from self-generated data; off-policy reuses data gathered by a different behavior policy.
Topic Knowledge Check
Exercise 1 of 3 • Test your architectural comprehension.
When a trained RL agent ships to production, which object actually runs at inference time?
How clear and actionable was this distributed systems breakdown?