Reinforcement learning from human feedback
RLHF uses human comparisons to learn a reward signal, then optimizes a model to produce responses that score better under that signal.
RLHF is a pipeline, not a single loss. A supervised model first produces candidate responses. People compare pairs, a reward model learns to predict those choices, and a reinforcement-learning algorithm updates the response policy against that learned score.
The reward model lets a small amount of human feedback guide a lot of training, but it is a stand-in for people, not people. In the summarization study, gentle training against the reward model gave better summaries. Pushing much further backfired: human raters liked the results less than the reward model’s scores implied. InstructGPT added a penalty for drifting away from the supervised model, to curb that over-optimizing.
An assistant study tracked that drift directly. Measured as KL divergence, it found reward rose roughly as a straight line in the square root of how far the policy had moved from its start. The labels also carry human choices. The researchers who write the labeling instructions shape what the model learns, along with the labelers themselves.
Some goals, like what makes a summary good, are hard to score without asking people.
Follow one prompt through RLHF: sample, rank, model, optimize.
- 1 · sampleHave the supervised model write two or more answers to the same prompt.
- 2 · rankAsk human labelers which answer in each pair they like better.
- 3 · modelTrain a reward model to predict those pairwise preferences.
- 4 · optimizeUse reinforcement learning to raise the predicted reward while keeping the policy close to where it started.
The policy learns from a model of human preferences, not from a person judging every update.
| Who | What they ask | What it works with |
|---|---|---|
| Assistant team | “Which answer is more helpful and harmless?” | Human rankings of response pairs |
| Reward modeller | “Does the scorer generalize to new prompts?” | Held-out preference accuracy |
| Alignment engineer | “Is reward rising because response quality improved?” | Human evaluations beside reward and KL drift |
- Comparing pairs lets people steer a system toward a goal without writing its reward function.
- A reward model trained on human comparisons can stand in as the reward during training.
- Researchers have used RLHF to train summarizers, instruction-following models, and chat assistants.
- Labelers do not always agree on which answer is better.
- Push hard enough on the reward model and its scores can stop matching what people actually prefer.
- A model trained this way can still make simple mistakes.
- It learns the tastes of its labelers and the people who wrote their instructions, not of every user.
Sources used
This explainer is written in original language. The links below support its factual claims.
- paperDeep reinforcement learning from human preferences, Christiano et al. · read 27 Sept 2026
- paperLearning to summarize from human feedback, Stiennon et al. · read 27 Sept 2026
- paperTraining language models to follow instructions with human feedback, Ouyang et al. · read 27 Sept 2026
- paperTraining a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback, Bai et al. · read 27 Sept 2026
- paperDirect Preference Optimization: Your Language Model is Secretly a Reward Model, Rafailov et al. · read 27 Sept 2026