TOPIC #257Advanced 16 min read

PRM vs ORM: Process vs Outcome Reward Models

AI
AI & ML Editorial
Report an issue
Key takeawayCore Concept Summary

A model writes a 12-step solution and gets one number: right or wrong. Which step caused the failure? That is credit assignment, and it is the crux of training reasoning. Outcome reward models grade the final answer; process reward models grade every step. This topic covers what each buys, what each costs, and how 2024-2026 practice blends them with rule-based verifiers.

01.The Problem: One Grade, Twelve Steps — Who Is at Fault?

A model solves a math problem in 12 steps. The final answer is wrong. Reward: 0.

Insight

Which of the twelve steps caused it?

The training signal says nothing. Step 7 had a sign error; steps 8-12 were actually careful, correct work built on the broken step. But a single "0" at the end punishes all twelve equally — and rewards them all equally when the answer luckily comes out right despite a broken path.

This is the credit assignment problem: deciding which part of a long trajectory deserves the praise or the blame.

It gets worse at reasoning scale. With a 30,000-token chain of thought and only a handful of decisive steps, outcome-only RL learns slowly — the optimizer has to average over many rollouts before discovering that step 7 is the culprit. And it can even increase wrong-but-fluent behaviour: fluent-but-wrong traces get reinforced whenever the final answer happens to be right by luck.

Two families address this:

  • ORM (outcome reward model / verifier): predicts correctness from the final answer or from the whole trace, without per-step judgments. Can be a learned classifier or, better, a rule: exact match, unit-test execution, proof-checker acceptance.
  • PRM (process reward model): scores each intermediate step (or token span), yielding dense signal for search and RL.

In one line each:

Insight

An ORM is the bouncer at the finish line: did you arrive? A PRM is the examiner standing at every step: was THAT right?

Everything below is the cost/benefit trade of putting examiners at every step.

Where the Reward Signal Lands 🎯

PRO Architecture Blueprint

Where the Reward Signal Lands 🎯

An ORM tells you whether the journey ended well. A PRM tells you which step was wrong - but only if its step labels are trustworthy, which is the whole difficulty.

Where the Reward Signal Lands 🎯
100%
Touchpad: Pinch to zoom • Drag to pan
Rendering visual architecture flowchart...
PRO & LIFETIME CURRICULUM

Unlock Topic #257: PRM vs ORM: Process vs Outcome Reward Models

You are viewing a preview. The full in-depth technical walkthrough, worked derivations, and code notebooks for this concept, along with self-assessment quizzes, are available with Pro or Lifetime Access.

Production Deep Dive

Failure modes, high-throughput bottlenecks, and real FAANG implementation decisions.

Interactive Blueprints

Interactive system topology diagrams, live parameter simulators, and downloadable SVG charts.

Knowledge Assessment

Staff-level multiple-choice quiz questions with instant feedback and answer explanations.

Cross-Device Progress Sync

Firebase Google authentication automatically syncs your completed topics and quiz scores.

Rate This Architecture ChapterFeedback & Rating

How clear and actionable was this distributed systems breakdown?