Preference Optimization
RLHF, DPO and RLAIF explained by what each one removes: how preference pairs teach judgement that demonstrations cannot, and why reward hacking and length bias are predictable.
Last updated
After this section you can
- Explain why ranking two answers teaches what writing one ideal answer cannot
- Place RLHF, DPO and RLAIF on the two axes that separate them: the optimiser and the labeller
- Diagnose reward hacking, length bias, sycophancy and over-refusal, including after DPO
Preference Optimization: RLHF, DPO & RLAIF
Supervised fine-tuning teaches a model to imitate one answer. Preference optimisation teaches it which of two answers is better, which is cheaper to label and captures qualities nobody can write down.
Ranking two answers is easier and more reliable than writing the ideal one, and it teaches judgement: length, tone, hedging, when to refuse. RLHF, DPO and RLAIF all learn from the same (prompt, chosen, rejected) pairs. They differ in who labels the pairs and how the model is optimised on them.