Reinforcement learning from AI feedback
RLAIF runs the usual RLHF training loop, but another AI model supplies some of the preference labels that people would normally give.
In standard RLHF, people rank model outputs, and those rankings drive the reinforcement learning step. RLAIF keeps that loop but hands the ranking job to an AI judge. The judge sees two candidate responses plus a principle from a constitution, and says which response is better. Its picks train a reward model, sometimes called a preference model. Choosing between two options is an older idea; early RLHF work had people compare pairs of short stretches of an agent’s behavior.
In the Constitutional AI paper, the model first critiqued and revised its own answers and was fine-tuned on the revisions. In its RL stage, harmlessness labels came from the AI, while helpfulness labels still came from people. Later experiments compared AI-labelled and human-labelled pipelines on summaries, helpful dialogue, and harmless dialogue.
People still write the rubric, so the list of principles shapes every label. How the judge is prompted changes how often it agrees with people; asking it to reason step by step tended to help. On summaries, the AI labeler agreed with human labels 78% of the time. For comparison, human labelers on a similar dataset agreed with each other only 73 to 77% of the time. The Lee et al. authors took this as a sign that RLAIF can stand in for human labels on these tasks. Human-feedback projects have collected fresh comparisons as often as every week to update their preference models and policies.
Good human preference labels cost a lot of money to gather.
Use a constitution to turn two responses into a training reward.
- 1 · sampleGenerate two candidate responses for each prompt.
- 2 · judgePrompt an AI model with evaluation principles and ask it to choose the better response.
- 3 · learnTrain a preference model on the resulting AI-labelled comparisons.
- 4 · optimizeUse that preference model as the reward signal for reinforcement learning.
RLAIF scales the act of applying a rubric; people still write the rubric.
| Who | What they ask | What it works with |
|---|---|---|
| Safety team | “Which response better follows the written principles?” | AI harmlessness labels mixed with human helpfulness labels |
| Post-training team | “Can feedback volume grow without equal growth in raters?” | Cost and agreement of AI labels |
| Evaluator | “Did the judge's biases carry into the policy?” | Human ratings of the trained policy |
- RLAIF can produce preference labels for under a tenth of what human labels cost.
- A constitution can make the principles used for feedback explicit.
- In tests on three text tasks, policies trained on AI labels improved about as much as policies trained on human labels.
- Biases that slip into AI labels can pass to the trained policy, which may then make them stronger.
- An AI judge can favour an answer just because of where it sits in the prompt.
- For high-stakes uses, the RLAIF authors still treat carefully trained human experts as the gold standard for labels.
Sources used
This explainer is written in original language. The links below support its factual claims.
- paperConstitutional AI: Harmlessness from AI Feedback, Bai et al. · read 27 Sept 2026
- paperRLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback, Lee et al. · read 27 Sept 2026
- paperTraining language models to follow instructions with human feedback, Ouyang et al. · read 27 Sept 2026
- paperTraining a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback, Bai et al. · read 27 Sept 2026
- paperDeep reinforcement learning from human preferences, Christiano et al. · read 27 Sept 2026