TOPIC #145Advanced 13 min read

RLHF: Reinforcement Learning from Human Feedback

AI
AI & ML Editorial
Report an issue
Key takeawayCore Concept Summary

SFT can copy ideal answers, but "helpful" is a ranking between many good answers, not one correct string. RLHF learns that ranking from human comparisons: fine-tune, collect preference votes, distill them into a reward model, then optimize with PPO on a KL leash. This is the pipeline that became ChatGPT — plus its costs, failure modes, and the cheaper alternatives it spawned.

The InstructGPT Three-Stage Pipeline 🔁

Human preferences are distilled into a scalar reward model, then PPO maximizes that reward under a KL leash to the SFT policy.

The InstructGPT Three-Stage Pipeline 🔁
100%
Touchpad: Pinch to zoom • Drag to pan
Rendering visual architecture flowchart...

01.The Problem: There Is No Single Correct Answer

You asked SFT (Topic 144: fine-tuning on (instruction, ideal-response) pairs) to teach a model to help.

It works — until two answers are both fine, and one is better:

  • Answer A: correct, but rambles for 600 words.
  • Answer B: correct, concise, cites its steps.

SFT says: "both are just text; copy whichever one was in the dataset."

It has no way to learn "B is better than A" — because better is not a token to predict. It is a preference.

So the question becomes:

Insight

How do you train a model when the supervision signal is a comparison, not an answer?

Reinforcement learning is the natural formalism for optimizing a scalar judgment of whole responses. RLHF is that idea wired to language models.

02.The Idea in Plain Words: Vote, Distill, Optimize

RLHF (Reinforcement Learning from Human Feedback) is:

Insight

A three-stage recipe that turns human "this is better" votes into a trainable score, then pushes the model toward higher scores without drifting off a leash.

The lineage, in plain words:

  • Christiano et al. (2017, arXiv 1706.03741) first showed deep RL policies could be trained from human comparisons instead of a hand-written reward function. Why it matters: humans reliably say "trajectory A is better" far more easily than they can write reward math on paper.
  • Ouyang et al. (2022, InstructGPT, arXiv 2203.02151) ported that recipe to language models.
  • The resulting product became ChatGPT in November 2022.

03.The Three Stages, Walked With Tiny Numbers

Stage 1 — SFT. Fine-tune the base model on demonstrations of ideal behavior (Topic 144). This defines the starting policy \pi_{ref} and the response distribution the reward model will see.

Stage 2 — Reward model. Prompt the SFT model with real user queries and sample ~4–9 responses per prompt. Contractors rank them. Train a classifier (usually the SFT model plus a scalar head) on those pairwise rankings with the Bradley–Terry loss (Topic 146).

Worked micro-example — prompt: "Explain gravity to a 10-year-old". Four sampled replies get ranked:

y2 (concise analogy)  ≻  y1 (correct but dry)  ≻  y4 (rambling)  ≻  y3 (wrong)

Every neighboring pair becomes one training row: (y2 beats y1), (y1 beats y4), (y4 beats y3)... The reward model learns numbers like r(y2)=+7, r(y1)=+4, r(y4)=−1, r(y3)=−8 consistent with those votes.

Stage 3 — PPO optimization. The policy generates responses; the reward model scores them; PPO updates the policy to raise reward, minus a per-token KL penalty \beta \cdot KL(\pi | \pi_{ref}) that prevents reward hacking and distribution collapse (Topic 147).

InstructGPT's headline: the 1.3B RLHF model was preferred over GPT-3 175B, and RLHF reduced harmful outputs — despite being trained on only ~33 iterations of ~10k comparison prompts. Alignment, it turned out, was surprisingly sample-efficient.

04.Visual Intuition: The Loop

code
   ┌──────────────────────────────────────────────┐
   │                                              │
   ▼                                              │
 [Policy model]── writes reply y ──► [Reward model] scores r(x,y)
   ▲                                        │
   │              reward = r − β·KL(leash:   │
   │                     "stay near SFT me") │
   └──── PPO nudges weights toward higher ◄──┘
              reward, one small step

Read the chain:

humans vote → votes train the reward model → reward model scores every new reply → PPO walks the policy uphill on those scores → leash keeps it from walking off a cliff.

One nuance the picture hides: humans vote only once, in stage 2. In stage 3, the reward model is a proxy doing millions of votes per day. That proxy assumption is where all the failure modes come from.

05.The Analogy: The Cooking Student and the Food Critic

Carry one story through all of it: a cooking student, a critic, and a leash.

  1. SFT: the student watches a master chef make ideal dishes and copies them. Good manners, fixed ceiling.
  2. Preference collection: the student now plates the same dish four ways; a human taster ranks the plates. Ranking plates is easy for a taster — writing a formula for "delicious" is impossible.
  3. Reward model: the student hires a sous-chef who watched all the tastings. The sous-chef now scores any plate 1–10 instantly. Cheap, tireless, fallible.
  4. PPO with KL leash: the student cooks thousands of plates, chasing sous-chef scores — but with a rule: "cook in the style you were trained in, don't go feral."

The failure modes name themselves: the student learns to please the sous-chef, not the original tasters — garnish spam, extreme salt, confidence theater. That is reward hacking. The leash (stay near your trained style) is what keeps the plates edible.

06.Failure Modes and Costs: Why Labs Built Alternatives

RLHF works, but it is the most operationally painful stage of post-training:

  • Engineering cost: PPO for LLMs runs four models simultaneously (policy, reference, value/critic, reward), needs on-policy rollouts (expensive generation inside training), and is notoriously unstable to hyperparameters. In our story: a student, a sous-chef, a style-keeper, and a kitchen judge, all on the payroll.
  • Proxy-objective risks: the reward model is a lossy compressed version of human judgment. Gao et al. (2023) measured first over-optimization (reward keeps rising while true quality falls) and later reverse-KL hacking where the policy finds text the RM likes but humans despise.
  • Bias transfer: contractor demographics, labeling guidelines, and fatigue are baked into the reward — sycophancy is the classic learned pathology (helpful ≠ true: agreeing with the user scores well).
  • Human data bottleneck: preference collection is slow and costly, motivating AI-feedback variants (RLAIF, Constitutional AI — Topic 149) and offline alternatives (DPO — Topic 148) that dominate 2023–2026 open-model training.

Nevertheless, frontier assistants in 2026 still use RLHF-derived pipelines — with the RL stage increasingly reserved for verifiable-reward training (math/code reasoning models à la o1, where the reward is a checker program, not a human proxy — the sous-chef replaced by a lab test).

Architectural Trade-offs & Production Realities

Architectural Advantages

  • Optimizes whole-response quality judgments SFT cannot express.
  • Sample-efficient alignment: tens of thousands of comparisons transformed behavior.
  • Proven lineage: the direct technology behind ChatGPT.

Trade-offs & Constraints

  • Four-model PPO infrastructure, unstable, generation-in-the-loop is slow and pricey.
  • Reward hacking against a proxy RM; over-optimization curves are easy to hit.
  • Encodes annotator bias (sycophancy, verbosity) as if it were ground truth.
Production Implementation in Big Tech
OpenAI• ChatGPT

ChatGPT is the productized InstructGPT recipe: GPT-3.5 base → SFT on demonstrations → reward model from human comparisons of sampled replies → PPO with KL penalty. OpenAI later added Instruct-RLHF variants and, from 2024, RL on verifiable rewards for reasoning models.

Staff+ Engineering Takeaways

  • RLHF = SFT policy + reward model from human pairwise rankings + PPO under a KL penalty to the reference policy.
  • It exists because "better/worse" over whole responses cannot be taught by imitation alone.
  • InstructGPT 1.3B beat GPT-3 175B on human preference: alignment is sample-efficient.
  • Core risks: reward hacking against a proxy RM, over-optimization, and baked-in annotator bias (sycophancy).
  • DPO (offline) and RLAIF/CAI (AI feedback) were direct responses to RLHF's cost and instability; RL now shines where rewards are verifiable.

Topic Knowledge Check

Exercise 1 of 3 • Test your architectural comprehension.

Exercise 1 of 30 answered
1

In RLHF for LLMs, the KL penalty between the policy and the reference model primarily prevents:

Rate This Architecture ChapterFeedback & Rating

How clear and actionable was this distributed systems breakdown?