RLHF: Reinforcement Learning from Human Feedback
SFT can copy ideal answers, but "helpful" is a ranking between many good answers, not one correct string. RLHF learns that ranking from human comparisons: fine-tune, collect preference votes, distill them into a reward model, then optimize with PPO on a KL leash. This is the pipeline that became ChatGPT — plus its costs, failure modes, and the cheaper alternatives it spawned.
The InstructGPT Three-Stage Pipeline 🔁
Human preferences are distilled into a scalar reward model, then PPO maximizes that reward under a KL leash to the SFT policy.
01.The Problem: There Is No Single Correct Answer
You asked SFT (Topic 144: fine-tuning on (instruction, ideal-response) pairs) to teach a model to help.
It works — until two answers are both fine, and one is better:
- Answer A: correct, but rambles for 600 words.
- Answer B: correct, concise, cites its steps.
SFT says: "both are just text; copy whichever one was in the dataset."
It has no way to learn "B is better than A" — because better is not a token to predict. It is a preference.
So the question becomes:
How do you train a model when the supervision signal is a comparison, not an answer?
Reinforcement learning is the natural formalism for optimizing a scalar judgment of whole responses. RLHF is that idea wired to language models.
02.The Idea in Plain Words: Vote, Distill, Optimize
RLHF (Reinforcement Learning from Human Feedback) is:
A three-stage recipe that turns human "this is better" votes into a trainable score, then pushes the model toward higher scores without drifting off a leash.
The lineage, in plain words:
- Christiano et al. (2017, arXiv 1706.03741) first showed deep RL policies could be trained from human comparisons instead of a hand-written reward function. Why it matters: humans reliably say "trajectory A is better" far more easily than they can write reward math on paper.
- Ouyang et al. (2022, InstructGPT, arXiv 2203.02151) ported that recipe to language models.
- The resulting product became ChatGPT in November 2022.
03.The Three Stages, Walked With Tiny Numbers
Stage 1 — SFT. Fine-tune the base model on demonstrations of ideal behavior (Topic 144). This defines the starting policy \pi_{ref} and the response distribution the reward model will see.
Stage 2 — Reward model. Prompt the SFT model with real user queries and sample ~4–9 responses per prompt. Contractors rank them. Train a classifier (usually the SFT model plus a scalar head) on those pairwise rankings with the Bradley–Terry loss (Topic 146).
Worked micro-example — prompt: "Explain gravity to a 10-year-old". Four sampled replies get ranked:
y2 (concise analogy) ≻ y1 (correct but dry) ≻ y4 (rambling) ≻ y3 (wrong)
Every neighboring pair becomes one training row: (y2 beats y1), (y1 beats y4), (y4 beats y3)... The reward model learns numbers like r(y2)=+7, r(y1)=+4, r(y4)=−1, r(y3)=−8 consistent with those votes.
Stage 3 — PPO optimization. The policy generates responses; the reward model scores them; PPO updates the policy to raise reward, minus a per-token KL penalty \beta \cdot KL(\pi | \pi_{ref}) that prevents reward hacking and distribution collapse (Topic 147).
InstructGPT's headline: the 1.3B RLHF model was preferred over GPT-3 175B, and RLHF reduced harmful outputs — despite being trained on only ~33 iterations of ~10k comparison prompts. Alignment, it turned out, was surprisingly sample-efficient.
04.Visual Intuition: The Loop
code┌──────────────────────────────────────────────┐ │ │ ▼ │ [Policy model]── writes reply y ──► [Reward model] scores r(x,y) ▲ │ │ reward = r − β·KL(leash: │ │ "stay near SFT me") │ └──── PPO nudges weights toward higher ◄──┘ reward, one small step
Read the chain:
humans vote → votes train the reward model → reward model scores every new reply → PPO walks the policy uphill on those scores → leash keeps it from walking off a cliff.
One nuance the picture hides: humans vote only once, in stage 2. In stage 3, the reward model is a proxy doing millions of votes per day. That proxy assumption is where all the failure modes come from.
05.The Analogy: The Cooking Student and the Food Critic
Carry one story through all of it: a cooking student, a critic, and a leash.
- SFT: the student watches a master chef make ideal dishes and copies them. Good manners, fixed ceiling.
- Preference collection: the student now plates the same dish four ways; a human taster ranks the plates. Ranking plates is easy for a taster — writing a formula for "delicious" is impossible.
- Reward model: the student hires a sous-chef who watched all the tastings. The sous-chef now scores any plate 1–10 instantly. Cheap, tireless, fallible.
- PPO with KL leash: the student cooks thousands of plates, chasing sous-chef scores — but with a rule: "cook in the style you were trained in, don't go feral."
The failure modes name themselves: the student learns to please the sous-chef, not the original tasters — garnish spam, extreme salt, confidence theater. That is reward hacking. The leash (stay near your trained style) is what keeps the plates edible.
06.Failure Modes and Costs: Why Labs Built Alternatives
RLHF works, but it is the most operationally painful stage of post-training:
- Engineering cost: PPO for LLMs runs four models simultaneously (policy, reference, value/critic, reward), needs on-policy rollouts (expensive generation inside training), and is notoriously unstable to hyperparameters. In our story: a student, a sous-chef, a style-keeper, and a kitchen judge, all on the payroll.
- Proxy-objective risks: the reward model is a lossy compressed version of human judgment. Gao et al. (2023) measured first over-optimization (reward keeps rising while true quality falls) and later reverse-KL hacking where the policy finds text the RM likes but humans despise.
- Bias transfer: contractor demographics, labeling guidelines, and fatigue are baked into the reward — sycophancy is the classic learned pathology (helpful ≠ true: agreeing with the user scores well).
- Human data bottleneck: preference collection is slow and costly, motivating AI-feedback variants (RLAIF, Constitutional AI — Topic 149) and offline alternatives (DPO — Topic 148) that dominate 2023–2026 open-model training.
Nevertheless, frontier assistants in 2026 still use RLHF-derived pipelines — with the RL stage increasingly reserved for verifiable-reward training (math/code reasoning models à la o1, where the reward is a checker program, not a human proxy — the sous-chef replaced by a lab test).
Architectural Trade-offs & Production Realities
Architectural Advantages
- Optimizes whole-response quality judgments SFT cannot express.
- Sample-efficient alignment: tens of thousands of comparisons transformed behavior.
- Proven lineage: the direct technology behind ChatGPT.
Trade-offs & Constraints
- Four-model PPO infrastructure, unstable, generation-in-the-loop is slow and pricey.
- Reward hacking against a proxy RM; over-optimization curves are easy to hit.
- Encodes annotator bias (sycophancy, verbosity) as if it were ground truth.
ChatGPT is the productized InstructGPT recipe: GPT-3.5 base → SFT on demonstrations → reward model from human comparisons of sampled replies → PPO with KL penalty. OpenAI later added Instruct-RLHF variants and, from 2024, RL on verifiable rewards for reasoning models.
Staff+ Engineering Takeaways
- RLHF = SFT policy + reward model from human pairwise rankings + PPO under a KL penalty to the reference policy.
- It exists because "better/worse" over whole responses cannot be taught by imitation alone.
- InstructGPT 1.3B beat GPT-3 175B on human preference: alignment is sample-efficient.
- Core risks: reward hacking against a proxy RM, over-optimization, and baked-in annotator bias (sycophancy).
- DPO (offline) and RLAIF/CAI (AI feedback) were direct responses to RLHF's cost and instability; RL now shines where rewards are verifiable.
Topic Knowledge Check
Exercise 1 of 3 • Test your architectural comprehension.
In RLHF for LLMs, the KL penalty between the policy and the reference model primarily prevents:
How clear and actionable was this distributed systems breakdown?