1:1 mentoring with Big Tech AI engineers
LLM & Agentic

Preference Optimization

RLHF, DPO and RLAIF explained by what each one removes: how preference pairs teach judgement that demonstrations cannot, and why reward hacking and length bias are predictable.

Last updated

After this section you can

  • Explain why ranking two answers teaches what writing one ideal answer cannot
  • Place RLHF, DPO and RLAIF on the two axes that separate them: the optimiser and the labeller
  • Diagnose reward hacking, length bias, sycophancy and over-refusal, including after DPO
48

Preference Optimization: RLHF, DPO & RLAIF

Supervised fine-tuning teaches a model to imitate one answer. Preference optimisation teaches it which of two answers is better, which is cheaper to label and captures qualities nobody can write down.

Key idea

Ranking two answers is easier and more reliable than writing the ideal one, and it teaches judgement: length, tone, hedging, when to refuse. RLHF, DPO and RLAIF all learn from the same (prompt, chosen, rejected) pairs. They differ in who labels the pairs and how the model is optimised on them.

Preference optimisation is two independent choices: who ranks, and how you optimise
Humans rank the pairs An AI judge ranks the pairs against written principles WHO LABELS → HOW IT OPTIMISES ↓ Reward model + RL train a scorer, then PPO with a KL leash Direct loss on pairs DPO: no reward model, no sampling loop RLHF InstructGPT, early chat models most machinery, most control RLAIF Constitutional AI’s RL stage cheap labels, auditable rules DPO on human pairs the usual choice outside big labs an ordinary training run DPO on AI-ranked pairs common for open models cheapest end to end Same data in every cell: (prompt, chosen, rejected). The labeller and the optimiser are two separate choices.

Related

More in LLM & Agentic

Get full access to all 74+ sections with code examples, diagrams, and interactive animations.

Unlock Premium