Post-training Beyond RLHF: RLAIF, Self-Play and Verifiable Rewards
Classic RLHF pays humans to rate answers — slow, expensive, fooled by long text, and useless for judging a 300-step proof. The 2024-2026 stack swapped in four alternatives: AI-labelled preferences (RLAIF/CAI), offline pair losses (DPO/KTO/SimPO), self-play, and programmatic verification (RLVR + GRPO). This topic is how each works and when to use which.
Four Post-training Families Past Classic RLHF 🛠️
Human preference data is expensive, slow, and weak on long reasoning. Labs replaced it with AI labels, offline pair objectives, self-play, and programmatic verification - often all four in one pipeline.
01.The Problem: Human Graders Cannot Scale
A base language model predicts text. It is not good at answering you until post-training teaches it.
The classic teacher was human preference data. In one sentence: humans rate pairs of answers, and the model learns to produce answers humans rate higher — that pipeline is RLHF (Reinforcement Learning from Human Feedback), the recipe behind InstructGPT and GPT-4.
It worked. Then it hit four walls, all of them structural:
- Annotation is slow and expensive, and quality depends on instructions, rubrics, and calibration; frontier models outgrow non-expert labelers on math, code, medicine, and long agentic traces. A labeler with 90 seconds per pair cannot referee a 30,000-token proof.
- Human preference is a proxy objective, so it is hackable: reward models learn to favour length, hedging, confident tone, and formatting ("reward hacking"/"style over substance" is the well-known failure documented across 2023-2025). The model learns to sound better, not be better.
- PPO on a chat model is infrastructure-heavy: actor, critic, reference, and reward models plus rollouts, with unstable KL control (KL = the leash keeping the tuned model near its starting point). Four models in memory at once, per training run.
- It suppresses exploration. On multi-step reasoning, human-labelled "good answers" cannot teach self-correction, backtracking, or 20-step tool chains. A human can mark the final answer good or bad; that teaches almost nothing about how to discover it.
DeepSeek-R1 (Jan 2025) is the clean public demonstration of the pivot: a base model RL-trained with rule-based verifiable rewards (exact answer match, format checks) plus rejection sampling on correct traces, with little or no human preference data for reasoning, produced long chains of thought containing self-reflection spontaneously. That reframed post-training as objective design, not annotation procurement.
So the question for 2024-2026 became:
Who — or what — replaces the human grader, and what exactly does it grade?
02.The Idea in Plain Words: Four Replacements for the Grader
Every method in this topic answers the same question — where does the "better/worse" signal come from? — with one of four replacements:
1. RLAIF/CAI: a model grades, following a written rubric. 2. Offline preference (DPO and kin): skip the grader; train directly on pairs of "this answer beat that one". 3. Self-play: the model is its own grader, opponent, and task-writer. 4. RLVR: a program grades — tests, calculators, proof checkers.
All of them produce the same underlying thing post-training needs: a signal saying which behaviour to repeat. They differ in four dimensions you should hold in your head like a table:
- Cost: programs are free per check; humans are the most expensive per judgment.
- Trust: a compiler cannot be flattered. A human grader cannot referee proofs. A judge model inherits its own biases.
- Coverage: programs only check what is mechanically checkable. Style, taste, and "was this helpful" need soft graders.
- Exploration: offline pair losses only teach "prefer A over B among these two existing samples". Online RL (PPO/GRPO with rollouts) lets the model discover new strategies by trying.
The classic RLHF stack for reference — three stages: SFT on human demonstrations → train a reward model on human pairwise preferences → optimize the policy with PPO under a KL penalty to a reference model. Each new family keeps one or two stages and deletes the rest. RLAIF swaps the humans in stage 2 for a model. DPO collapses stages 2-3 into a single loss. RLVR deletes the reward model entirely and installs a program.
03.A Simple Worked Example: One Math Prompt, Four Regimes
Prompt: "A shop marks a 90-euro jacket down 20%, then adds 19% tax. Final price?" Correct: 90 × 0.8 × 1.19 = 85.68.
Sample four model answers and grade them under each regime:
- A: "… 72 × 1.19 = 85.68. So 85.68 euros."
- B: "85.68. (short)"
- C: "It depends on many factors, but roughly 86 euros is likely."
- D: "I first compute 90 − 20 = 70… final price 83.30."
Regime 1 — RLHF (human raters). People tend to prefer A: complete, confident, shows work. C might beat D. Fine — but you paid 4 labelers per pair, and now imagine the proof is 200 steps long.
Regime 2 — RLAIF (judge model with a rubric). A judge scores each answer 1-5 on correctness + clarity, same result as the humans but free at scale — and notice the trap: the judge also tends to like A over B, feeding the length bias, because "shows work" correlates with quality in its training.
Regime 3 — DPO (offline pairs). No scores at all. Build pairs: (A beats D), (A beats C). The loss pushes up log-probability of the chosen answer relative to the rejected. One model, one gradient step, no reward model.
Regime 4 — RLVR + GRPO (program grader). Check each answer against the computed value: rewards = A 1, B 1, C 0, D 0. Now GRPO turns these four into learning signal without any critic network:
codegroup of 4 samples per prompt rewards: [1, 1, 0, 0] mean = 0.5 advantage: [+, +, −, −] ← each reward minus the group mean (÷ spread, so 1-vs-0 gaps become usable step sizes) update: raise probability of A and B's tokens, lower C and D's
The model discovered nothing new here — but if 300 of 1,000 rollouts found a novel valid route to 85.68, the checker rewards it too. That is RLVR's magic: exploration paid for by an answer key, not an opinion.
04.Visual Intuition + the Analogy: The Grading Economy
Carry one analogy through the rest: a school building its exam-grading system.
- RLHF = hire 10,000 part-time human graders. Works. Gets slow. Graders get tired, favour long essays with tidy handwriting, and cannot grade a doctoral thesis in quantum mechanics. (That last part is the frontier-model regime: labelers are out of their depth.)
- RLAIF / Constitutional AI = train one grader-bot on a written rubric and let it grade everything. Tireless, consistent, cheap per grade. But it loves essays that look like its rubric — including wrong ones with confident formatting. And its rubric is auditable: you can read exactly what "good" means, version it, and blame it.
- DPO and offline kin = skip numeric grades entirely. Hand the tutor two scripts per question — "this one is better" — and change the student's habits directly. No grading office (reward model), no exam hall with new attempts (rollouts). The cheapest operation in the school.
- Self-play = the student writes their own practice questions, answers them, and grades them. Explosive if the answer key is real. A closed loop of self-delusion if the student is also the answer key.
- RLVR = only grade questions with answers at the back of the book. Objective, infinite, unhackable — but the back of the book only exists for math, code, and logic, not for "write a moving poem".
The map of the field:
codecan it be checked by a PROGRAM? no yes need new ┌─────────────────┐ ┌─────────────────┐ behaviours │ online RL: │ │ RLVR + GRPO │ ← discovery (exploration) │ RLHF / RLAIF │ │ (PPO w/ rules) │ works here ├─────────────────┤ ├─────────────────┤ happy with │ offline: │ │ rejection-sampling better │ DPO / KTO / │ │ SFT: keep only │ ranking of │ SimPO / ORPO │ │ correct traces │ samples │ (+ self-play) │ │ │ └─────────────────┘ └─────────────────┘ cost: humans ≫ judges ≫ pairs programs ≈ free trust: humans & judges are hackable; programs are not
Real pipelines are hybrids: usually offline pairs first (cheap alignment of what the model already says), then RLVR where an answer key exists, then a thin rubric-bot (CAI) pass for style and refusal behaviour.
05.RLAIF and Constitutional AI: AI as the Labeler
RLAIF (Anthropic 2023 line, then widely adopted at Google-Scale in 2024) replaces the human preference dataset with model-generated pairwise labels or LLM-judged scores, then proceeds with the usual RM plus PPO. Reported results across 2024 papers are consistent: on helpfulness tasks AI labels match human-supervised performance; on harmlessness they often exceed it, because the labelling model applies rubrics more consistently than fatigue-prone humans.
Constitutional AI / RLAIF-C adds an explicit principle set (a written "constitution": e.g. "decline requests for weapons instructions, but answer the underlying science"): the model critiques its own output against the constitution, revises it, and the revised pairs become training data (supervised) or preference pairs (RL). This is the operational template most labs reused in 2025-2026 for safety behaviour, because principles are auditable and versioned while human annotation is not.
Practical engineering notes:
- Labeler bias is inherited: the labelling model's verbosity, sycophancy (agreeing with the user to please them), and format preferences propagate. Mitigate with swapped-order judging (judge the pair both ways; contradictions get dropped), rubric-conditioned judging, length penalties, and periodic human arbitration on a stratified audit set.
- Best-value target is pairwise labels with a cheap labeler plus disagreement sampling — uncertain pairs (where the judge flip-flops) are where annotation budget pays, because easy pairs teach nothing new.
- Keep labeler prompt, decoding parameters, and model version pinned; changing the labeler mid-run shifts the reward distribution under your optimizer. (Swap the rubric-bot mid-semester and last month's grades stop being comparable to this month's.)
06.Offline Preference Objectives: DPO and Its Successors
The 2023-2025 simplification wave removed the reward model and online rollouts — the grading office and the exam hall both:
- DPO: reparameterize the RLHF objective as a classification loss on preference pairs, using the implicit reward log-ratio between policy and reference. In plain words: instead of training a grader and then optimizing, directly nudge "chosen" answers up and "rejected" answers down, with the reference model as the anchor so the nudges do not run away. Cheap, stable, easy to ship; sensitive to reference drift and to pairs where both answers are bad (a "preference" between two wrong poems is noise).
- IPO / ORPO / KTO: variants addressing over-optimization (IPO), dropping the reference model entirely (ORPO), or working from binary good/bad labels instead of pairs (KTO — convenient because your production thumbs-up/down data already has that shape: users click a button, they never compare two answers).
- SimPO (2024): reference-free, length-normalized reward with a margin; strong results without a KL anchor, and it directly attacks the length bias problem — because the score is per token, a longer answer no longer wins by accumulation.
The 2024-2025 debate ("is DPO actually better than well-tuned PPO?") resolved pragmatically: offline methods win on cost, reproducibility, and iteration speed; online RL wins when you need exploration, long trajectories, or on-policy correction. Hence the standard hybrid in 2026: SFT → offline preference/iterative DPO on model-generated pairs → RLVR with PPO or GRPO for verifiable skills → CAI-style safety pass.
The iterative version — the model generates its own pairs, the judge labels them, and offline training closes each round — is the common 2025 pattern:
import torch, torch.nn.functional as F
def simpo_loss(pi_logp, ref_free_logp, chosen_idx, rejected_idx, beta=2.0, gamma=0.5):
# length-normalized implicit rewards, no reference model needed
rc = (pi_logp.gather(1, chosen_idx[:, None]) / chosen_idx.size(1)).squeeze(1)
rr = (pi_logp.gather(1, rejected_idx[:, None]) / rejected_idx.size(1)).squeeze(1)
return -F.logsigmoid(beta * (rc - rr) - gamma).mean()
for round_i in range(3):
prompts = sample_prompt_bank(bank, n=200_000)
pairs = []
for p in prompts:
a, b = (policy.generate(p, temperature=1.0) for _ in range(2)) # on-policy samples
win = judge.preference(p, a, b, swap=True) # RLAIF label, order-robust
pairs.append((p, a, b) if win == "a" else (p, b, a))
policy = offline_pref_train(policy, pairs, objective="simpo")
# keep a small human-audited slice to catch judge drift07.Self-Play, Self-Rewarding, and Zero-Data Bootstrapping
The student writing their own exams, four ways:
- SPIN (self-play fine-tuning, 2024): treat human SFT data as "real" and model output as "fake"; the current model is the discriminator, and the loss teaches the policy to be indistinguishable from (and better than) the reference responses. Turns alignment into a two-player game without new labels.
- Self-Rewarding Language Models (2024): the model judges its own outputs via LLM-as-judge prompts, producing its own reward signal, then applies RL (DPO in the original work). Capability and alignment improve together across iterations — but the ceiling is the judge's own bias, and reward drift can accumulate. The student is grading their own exam with their own notes.
- Adversarial/self-generated curricula: an opponent model proposes progressively harder prompts the policy fails on (Kimi k1.5-style generative-reward-model adversarial prompting in 2025), which is how you scale past hand-written task banks. The school adds a question-writer whose only job is to make questions that just defeat the current student.
- Zero/absolute-zero style self-play (2025): the model proposes tasks it can verify (code with tests, self-contained problems with checkable answers) and trains on them without any human data, using proposer/solver separation so the task generator is rewarded for producing tasks that are just hard enough. The answer key is saved before the student learns to write the exams.
- Agentic self-improvement of scaffolding: Darwin Gödel Machine (2025) showed populations of agents rewriting their own coding-agent scaffolds with verification and archive-based selection — self-improvement at the system level rather than the weight level.
Where self-play breaks: no external ground truth means the loop converges to the judges' shared priors (mode collapse, style homogenization, and sometimes confident consensus on wrong answers). Rule-based verifiers, held-out human audits, or a stronger external teacher are the standard escape hatches. (The student-and-answer-key-by-the-same-hand failure: everyone in the class agrees the wrong answer is right.)
08.RLVR, GRPO, and a 2026 Post-training Plan
RLVR (reinforcement learning with verifiable rewards). Reward comes from a program: unit tests, string/math match, formal proof checkers, SQL execution, structured-schema validity. Because the reward is objective, you can run huge rollout counts and long chains without a learned reward model, and over-optimization is far harder. The cost: verifiable domains are narrow (math, code, logic, extraction), so style and safety still need other objectives.
GRPO (group-relative policy optimization). For each prompt, sample G completions; score them; set each token's advantage to the (standardized) deviation of its sequence reward from the group mean, so the critic/value network disappears. In 2024-2026 GRPO became the de facto open-frontier choice (DeepSeek Math/R1, Qwen3, Kimi k1.5 lineage) because it cuts memory and instability roughly in half versus PPO at equal rollout budget, and clips KL explicitly instead of via a critic. (You already saw the mechanics in the worked example: [1,1,0,0] → +,+ against −,− around the group mean.)
What practitioners add on top:
- Curriculum/difficulty filtering: prompts with a pass rate near 0 or 1 give ~zero gradient (all samples get the same reward → no advantage to learn from); sample the middle — the "just hard enough" band.
- Reward shaping for format plus correctness, with heavy penalties for verifier-gaming (test edits, answer leakage, output truncation tricks).
- Length control, since RLVR rewards reasoning length indirectly and models over-think.
- Rejection-sampling SFT interleaves with RL: distil successful traces back into the policy to stabilize.
Putting it together for a mid-size model (say 7-70B) targeting agentic product quality:
- SFT on distilled, verifier-filtered traces (shortest correct answer wins), including error-recovery turns.
- Offline preference (SimPO/DPO) on self-generated pairs labelled by an order-swapped judge plus a human-audited calibration set; use KTO where you only have thumbs data.
- RLVR with GRPO on verifiable capability slices (code with hidden tests, structured-output validity, retrieval-answerable QA), difficulty-filtered.
- CAI-style safety pass against a versioned principle set, with red-team prompts from an adversarial generator.
- Continuous judge calibration — the moment you scale AI labels, the judge is your reward model, and reward models need monitoring for hacking, drift, and style bias.
Instrument each stage with win-rate deltas against the previous checkpoint, length distributions, refusal/false-refusal rates, format-validity rates, and a fixed human-eval slice so you notice when the automated signal and the real signal diverge.
Architectural Trade-offs & Production Realities
Architectural Advantages
- AI labels and programmatic rewards scale: annotation no longer caps data volume, capability, or iteration speed.
- Offline objectives (DPO/KTO/SimPO/ORPO) remove the reward model and rollouts: far cheaper, reproducible, easy to ablate.
- RLVR plus GRPO delivers objective rewards and long-horizon reasoning gains that human preference cannot teach.
- Self-play bootstraps data where none exists (new language, new tool, new domain) and can target just-hard-enough tasks.
- Constitutional/rubric-based safety data is auditable and versioned, unlike anonymous human labeler judgments.
Trade-offs & Constraints
- Judge bias is your new reward bias: verbosity, sycophancy, format worship, and self-preference propagate silently.
- Self-play can converge on confident shared errors and homogenize outputs; no external ground truth means no ceiling detection.
- Verifiable rewards cover a narrow slice of what users want; style, taste, and open-ended help still need soft objectives.
- GRPO/PPO still need rollout infrastructure, careful KL and length control, and curriculum tuning; naive setups plateau or collapse.
- Reward hacking gets more sophisticated as verification gets more creative (test edits, answer leakage, verifier-specific quirks).
DeepSeek-R1 used large-scale GRPO over prompts whose answers are machine-checkable (exact-match math, code tests), combined with cold-start SFT on filtered correct traces and a final RLAIF-style preference phase for general-purpose helpfulness; the result was models that emit long self-correcting chains without human-labelled reasoning data. Qwen3 applied the same idea at production scale, switching between thinking and non-thinking modes and using RLVR plus distilled traces, then shipping distilled students that inherit the reasoning behaviour.
Staff+ Engineering Takeaways
- Classic RLHF (SFT → human-preference RM → PPO) is being demoted to a thin alignment layer because human labels cap scale, misjudge long reasoning, and are gameable.
- RLAIF/Constitutional AI substitute rubric-following model labels and self-critique-and-revise, giving auditable, versionable alignment data.
- Offline preference objectives (DPO, IPO, ORPO, KTO, SimPO) drop the reward model and rollouts: cheap and reproducible, weaker when exploration is needed.
- Self-play (SPIN, self-rewarding, zero-data task proposal, scaffold self-modification) bootstraps data but needs external grounding to avoid converging on shared errors.
- RLVR with GRPO is the 2024-2026 default for verifiable skills: objective rewards, huge rollouts, no critic; DeepSeek-R1 is the canonical public demonstration.
- Whichever labeler you use becomes your reward model, so judge calibration and reward-hacking monitoring are permanent responsibilities.
Topic Knowledge Check
Exercise 1 of 3 • Test your architectural comprehension.
What does GRPO eliminate relative to PPO-based RLHF?
How clear and actionable was this distributed systems breakdown?