PRM vs ORM: Process vs Outcome Reward Models
A model writes a 12-step solution and gets one number: right or wrong. Which step caused the failure? That is credit assignment, and it is the crux of training reasoning. Outcome reward models grade the final answer; process reward models grade every step. This topic covers what each buys, what each costs, and how 2024-2026 practice blends them with rule-based verifiers.
01.The Problem: One Grade, Twelve Steps — Who Is at Fault?
A model solves a math problem in 12 steps. The final answer is wrong. Reward: 0.
Which of the twelve steps caused it?
The training signal says nothing. Step 7 had a sign error; steps 8-12 were actually careful, correct work built on the broken step. But a single "0" at the end punishes all twelve equally — and rewards them all equally when the answer luckily comes out right despite a broken path.
This is the credit assignment problem: deciding which part of a long trajectory deserves the praise or the blame.
It gets worse at reasoning scale. With a 30,000-token chain of thought and only a handful of decisive steps, outcome-only RL learns slowly — the optimizer has to average over many rollouts before discovering that step 7 is the culprit. And it can even increase wrong-but-fluent behaviour: fluent-but-wrong traces get reinforced whenever the final answer happens to be right by luck.
Two families address this:
- ORM (outcome reward model / verifier): predicts correctness from the final answer or from the whole trace, without per-step judgments. Can be a learned classifier or, better, a rule: exact match, unit-test execution, proof-checker acceptance.
- PRM (process reward model): scores each intermediate step (or token span), yielding dense signal for search and RL.
In one line each:
An ORM is the bouncer at the finish line: did you arrive? A PRM is the examiner standing at every step: was THAT right?
Everything below is the cost/benefit trade of putting examiners at every step.
Where the Reward Signal Lands 🎯
Where the Reward Signal Lands 🎯
An ORM tells you whether the journey ended well. A PRM tells you which step was wrong - but only if its step labels are trustworthy, which is the whole difficulty.
Unlock Topic #257: PRM vs ORM: Process vs Outcome Reward Models
You are viewing a preview. The full in-depth technical walkthrough, worked derivations, and code notebooks for this concept, along with self-assessment quizzes, are available with Pro or Lifetime Access.
Failure modes, high-throughput bottlenecks, and real FAANG implementation decisions.
Interactive system topology diagrams, live parameter simulators, and downloadable SVG charts.
Staff-level multiple-choice quiz questions with instant feedback and answer explanations.
Firebase Google authentication automatically syncs your completed topics and quiz scores.
How clear and actionable was this distributed systems breakdown?