TOPIC #198Advanced 11 min read

Value Function and Q-Function

AI
AI & ML Editorial
Report an issue
Key takeawayCore Concept Summary

V(s) is the fair price of standing in a state — the long-run reward you can expect from there. Q(s,a) is the price of standing there AND making a specific move. That second number is decision-ready: with Q you can act without any model of the world, and subtracting V from Q gives the advantage that powers modern policy methods.

V(s) vs Q(s,a) Backup 📊

V(s) scores how good it is to BE in a state under policy pi. Q(s,a) scores how good it is to TAKE action a and then follow pi. They relate as V(s) = max_a Q(s,a) under an optimal policy, and Q feeds back into V via the Bellman optimality backup.

V(s) vs Q(s,a) Backup 📊
100%
Touchpad: Pinch to zoom • Drag to pan
Rendering visual architecture flowchart...

01.The Problem: A Cheat Sheet Is Worthless Until You Score It

Where we are in the story: topic 196 gave us the board game (the MDP). Topic 197 gave us the player's rule card (the policy).

Insight

How do you know whether a rule card is any good? And how do you find a better one?

You need numbers. You need to answer questions like:

  • "I just landed on this square — am I in a good spot or a doomed one?"
  • "From this square, is it better to go North or East?"
  • "By how much, though? I want the margin, not just a ranking."

Reinforcement learning reduces all three questions to one prediction problem: estimate the long-run reward an agent can expect from each situation. Two functions do that job — the state-value V and the action-value Q — and their difference, the advantage, is the workhorse of every modern method.

Sutton's one-line distinction, which we will earn by the end of this topic:

Insight

V(s) is how good it is to BE in a state. Q(s,a) is how good it is to DO a thing there.

02.The State-Value Function V_pi(s): The Fair Price of Standing Here

The state-value function answers the question:

Insight

Starting from state s, and then following policy π forever, what return do I expect?

(Recall the return from topic 196: G_t = R_{t+1} + γR_{t+2} + γ²R_{t+3} + …, the discounted sum of future rewards.)

Written out:

v_π(s) = E_π[G_t | S_t = s] = E_π[Σ_{k=0}^{∞} γᵏ R_{t+k+1} | S_t = s]

Unpack the notation:

  • E_π[…] means expectation (average) — and the subscript π says over which futures: the die of the environment rolls, and the policy rolls its own dice, and we average over every possible outcome of both.
  • | S_t = s means "given that we start here."
  • So v_π(s) is a single number: the fair price of being in state s, assuming you play by π from now on.

The phrase "assuming you play by π" is the whole point. Value is explicitly conditioned on a policy. The same square has different values for different players:

  • A chess position worth +2 for a grandmaster (whose policy punishes opponent mistakes) is nearly worthless for a beginner whose policy misses the tactic entirely.
  • The hallway square right before the exit is valuable to an agent whose policy heads for doors — and valueless to one whose policy wanders away.

Change π, and every number in the value table changes. A value on its own, with no π attached, is meaningless.

03.The Action-Value Function Q_pi(s,a): The Price of Making This Move

The action-value function — the famous Q-function — answers the more useful question:

Insight

If I am in state s, take action a right now (no matter what the policy says), and then follow π, what return do I expect?

q_π(s,a) = E_π[G_t | S_t = s, A_t = a]

Read the formula as a two-act story: act one is forced ("take a, whatever it costs"), act two is policy-driven ("then keep following π"). Q is the expected payoff of that plan.

Why is Q more powerful than V? Because choosing an action from Q requires no model of the world. If you know Q*(s,a), the optimal policy simply falls out:

π*(s) = argmax_a Q*(s,a)

No dice, no planning, just a lookup-and-pick-best. By contrast, if you only know V(s), picking the best action would additionally require the transition model P(s'|s,a) — you would have to simulate each action forward to see which destination has the higher value.

That contrast is the entire reason Q-learning and DQN learn Q, not V: Q is model-free and decision-ready; V is model-free but action-blind.

04.A Simple Worked Example: One State, Three Actions

Numbers. At some state s, under policy π, the three possible actions have action-values:

Q(s,a0) = 7.0, Q(s,a1) = 4.0, Q(s,a2) = 9.0

and the policy's dice at this state are π(a0|s) = 0.5, π(a1|s) = 0.3, π(a2|s) = 0.2.

Step 1 — get V from Q. Value of the state is the policy-weighted average of action values:

V(s) = 0.5×7.0 + 0.3×4.0 + 0.2×9.0 = 3.5 + 1.2 + 1.8 = 6.5

Step 2 — act greedily. The best action here is a2 (Q = 9.0). If π were playing otherwise, the improvement theorem from topic 197 says rewriting this line to "take a2" can only help.

Step 3 — compute advantages. How much better or worse than the state's average is each action?

A(s,a) = Q(s,a) − V(s)

  • A(s,a0) = 7.0 − 6.5 = +0.5 — slightly better than average.
  • A(s,a1) = 4.0 − 6.5 = −2.5 — clearly bad; the policy wastes 30% of its dice here.
  • A(s,a2) = 9.0 − 6.5 = +2.5 — the star move.

Notice why the advantage matters: the state itself scores 6.5, which is "good." But blindly rewarding everything at a good state would credit a1 for a situation it did nothing to create. Subtracting V(s) removes the situation's contribution and leaves only the action's contribution. The code below does exactly these three steps.

python— Relating V, Q, and the advantage for a single state — the worked example, executable
import numpy as np

# q(s, .) for one state under policy pi, and the policy probabilities
q_pi = np.array([7.0, 4.0, 9.0])     # Q(s, a0..a2)
pi   = np.array([0.5, 0.3, 0.2])     # pi(a|s)

V = (pi * q_pi).sum()                # V(s) = sum_a pi(a|s) Q(s,a)
advantage = q_pi - V                 # A(s,a) = Q(s,a) - V(s)

print(f"V(s)={V:.2f}")               # 6.50
print(f"A(s,.)={advantage.round(2)}") # [+0.5 -2.5 +2.5]
print("greedy action:", int(np.argmax(q_pi)))  # a2

05.Visual Intuition: Values Flow Backward From the Goal

Picture a corridor of squares, like the mermaid diagram above. The goal pays the most, and value seeps backward from it — each square's price is "what I collect now, plus the discounted price of where I land" (that sentence is the Bellman equation, topic 199, which you will meet next).

code
  corridor toward the goal (values from the diagram):

  s0 ──────► s1 ──────► s2 ──────► goal
  V=6.6      V=6.9      V=7.1      V=7.5
    │          │
    │ right:   │ right:
    │ Q=7.1    │ Q=8.0     ← Q: the price of each specific move
    │ left:    │
    │ Q=4.2    │
    ▼
  Q=2.0 dead-end            ← the move that wrecks the plan
  • Each square gets one number: V.
  • Each decision gets its own number: Q. Same square, different moves, very different prices — at s0, "right" is worth 7.1 while the dead-end left is worth 2.0.
  • V(s) sits between the Q's of its actions: it is their weighted average under π (or their maximum, under the optimal policy).

Interview-worthy one-liner: V grades the place, Q grades the move from the place, and the advantage grades the move relative to the place's grade.

06.Optimal Value Functions: The Scorecard of the Best Possible Player

Among all policies, there exist optimal value functions that dominate every other policy in every state:

V*(s) = max_π v_π(s), Q*(s,a) = max_π q_π(s,a)

Two facts to lock in:

  1. A policy that is greedy with respect to Q* is guaranteed optimal — π*(s) = argmax_a Q*(s,a) — even though you never explicitly compared policies.
  2. Nobody hands you V* or Q*. The central practical goal of temporal-difference and deep RL methods (topics 200–202) is to estimate Q* from raw experience — samples of play — rather than solving the MDP analytically.

So the algorithm roadmap for the rest of this phase writes itself: find a way to learn the scorecard (Q*) without knowing the rules of the game (the model P), then ship the argmax cheat sheet.

07.The Advantage Function: Why Modern Methods Subtract V From Q

Modern policy-gradient methods (A2C, A3C, PPO, GAE — topics 203–205) rarely use raw Q. They use the advantage:

A_π(s,a) = Q_π(s,a) − V_π(s)

One sentence: the advantage says how much better action a is than the average action at s.

Why bother subtracting? Variance. Remember that value is an expectation over lucky and unlucky futures. If you nudge the policy using raw returns, a state that is simply good (high V) makes every action taken there look great — the gradient credits the action for the situation it happened to be standing in. Subtracting the baseline V(s) strips the situation out and keeps only the action's own contribution. That single trick massively reduces gradient variance without biasing it, and it is one of the most important engineering insights in deep RL.

In practice:

  • The critic network in actor-critic methods is a learned V_π, used exactly as this baseline (topic 204).
  • A(s,a) = Q − V estimated with TD errors becomes GAE, the advantage estimator inside PPO, which is inside RLHF — the machinery that steers ChatGPT-class models (topic 203 shows the math).
  • Read the sign of an advantage as coaching: positive → "do this more," negative → "do this less."

Architectural Trade-offs & Production Realities

Architectural Advantages

  • Q makes optimal action selection model-free via a simple argmax.
  • V provides a low-dimensional baseline that de-noises policy gradients.
  • Bootstrapping V/Q yields low-variance, sample-efficient TD targets (topic 200).

Trade-offs & Constraints

  • Value is always policy-dependent — V(s) is meaningless without a pi attached.
  • Q(s,a) grows with the action space, so it is hard to use for huge or continuous action sets.
  • Bootstrapped values accumulate their own estimation error (compounding bias).
Production Implementation in Big Tech
Waymo / autonomous driving critics• Q-values for maneuver evaluation

A learned Q(s,a) scores long-term safety/comfort of candidate maneuvers (lane change, brake, yield) from a rich perception state, letting a planner pick argmax_a Q rather than hand-tuning a cost for every action.

Staff+ Engineering Takeaways

  • V_pi(s) = expected discounted return from state s, conditional on then following policy pi — fair price of the place.
  • Q_pi(s,a) = expected return after taking a in s then following pi — decision-ready: pi*(s) = argmax_a Q*(s,a) needs no model.
  • V(s) = sum_a pi(a|s) Q(s,a); under optimality V*(s) = max_a Q*(s,a).
  • Values flow backward from the goal; the Bellman equation (next topic) makes that recursion exact.
  • The advantage A(s,a) = Q(s,a) - V(s) removes the situation-level baseline, cutting policy-gradient variance without adding bias.

Topic Knowledge Check

Exercise 1 of 3 • Test your architectural comprehension.

Exercise 1 of 30 answered
1

Given an optimal Q-function, how do you get the optimal policy without any model of the world?

Rate This Architecture ChapterFeedback & Rating

How clear and actionable was this distributed systems breakdown?