Direct Preference Optimization
DPO teaches a language model to prefer the better of two answers by learning from the comparison itself, with no separate reward model trained first.
DPO, short for direct preference optimization, is a way to teach a language model which of two answers people like better. The older route, RLHF, has two stages, training a reward model to imitate human judgments and then using reinforcement learning to tune the language model against that reward model. The DPO authors call that pipeline complex and often unstable. DPO skips the separate reward model and trains on the comparisons with one simple classification-style loss.
Each training record holds a prompt, a chosen answer and a rejected answer. In this setting the “policy” just means the model being trained. A log probability is a model’s confidence that it would write a given answer, on a log scale where higher means more likely. DPO asks two models for these numbers. One is the policy, and the other is a reference model that only supplies comparison scores. In Hugging Face’s TRL library the reference defaults to the model as it was before DPO training began. Because it stays fixed, its scores can even be worked out once for the whole dataset before training starts.
For each model, subtract the rejected answer’s log probability from the chosen answer’s to get a gap. Training pushes the policy’s gap wider than the reference’s gap, with no reward model anywhere in the loop. A setting called beta decides how far the policy may drift from the reference, and a higher beta keeps it closer. One surprise is that the gap usually widens because the rejected answer becomes less likely, not because the chosen one becomes more likely.
Simpler does not mean assumption-free. Azar and colleagues note that RLHF assumes a preference between two answers can be swapped for a separate score on each answer. They point out that DPO still leans heavily on that assumption. The data can also go stale, since preference sets are normally gathered once, up front, and left alone while the model changes. Online variants fix this by sampling two new answers from the current model at every step and asking another model to pick the better one. Other losses compete too, and the KTO authors report matching or beating preference-based methods while learning only whether each single answer is good or bad.
Standard RLHF puts a reward model and a reinforcement-learning loop between the preference data and the final model.
Turn one preference pair into a direct update.
- 1 · pairStore each prompt with one chosen answer and one rejected answer.
- 2 · scoreAsk the trainable policy and the fixed reference model how likely each answer is, as log probabilities.
- 3 · compareTurn into a loss how much more the policy favours the chosen answer than the reference does.
- 4 · updateChange the policy's weights to lower that loss while the reference stays as it was.
DPO removes the reward-model stage; it still needs preference pairs and a reference model.
| Who | What they ask | What it works with |
|---|---|---|
| Post-training team | “Can we learn from ranked answers without running PPO?” | DPO loss and held-out preference evaluations |
| Data curator | “Are chosen and rejected answers truly distinguishable?” | Pair quality and annotator agreement |
| Researcher | “Is the offline preference set stale for the current policy?” | Offline versus online preference results |
- It drops the separate reward model and the reinforcement-learning loop from basic preference training.
- Training runs on one ordinary classification-style loss over preference pairs.
- Libraries such as Hugging Face TRL train it from records with prompt, chosen and rejected fields.
- Noisy or wrong preference labels still hurt, which is why variants such as Robust DPO exist.
- Pairs collected ahead of time drift out of step with the model as it trains.
- No single preference loss is best everywhere, and the right choice depends on the setting.
Sources used
This explainer is written in original language. The links below support its factual claims.
- paperDirect Preference Optimization: Your Language Model is Secretly a Reward Model, Rafailov et al. · read 27 Sept 2026
- docsDPO Trainer, Hugging Face · read 27 Sept 2026
- paperA General Theoretical Paradigm to Understand Learning from Human Preferences, Azar et al. · read 27 Sept 2026
- paperDirect Language Model Alignment from Online AI Feedback, Guo et al. · read 27 Sept 2026
- paperKTO: Model Alignment as Prospect Theoretic Optimization, Ethayarajh et al. · read 27 Sept 2026