Constitutional AI
Constitutional AI trains and evaluates model behavior against an explicit written set of principles.
Constitutional AI makes a set of behavioral principles an explicit part of post-training. In its original form, a model first generates a response, critiques that response under a sampled principle, and revises it. Those revisions become supervised fine-tuning data.
In a second phase, the fine-tuned model writes two answers to the same prompt. An AI judge picks the better one, guided by a principle from the constitution. A preference model learns from many of those picks, mixed with human labels for helpfulness, and reinforcement learning then uses it to score the assistant. The paper calls this reinforcement learning from AI feedback, or RLAIF. A later study by a different team tried RLAIF on three jobs: summarizing text, giving helpful chat answers, and giving harmless ones. On all three, it did about as well as training on human labels (RLHF).
The constitution makes intended behavior easier to inspect and revise, but it does not mechanically determine outputs. The model has to interpret the words and carry them over to cases the authors did not foresee, and even Anthropic expects parts of Claude’s behavior to fall short of its own constitution. The original authors also flag risks. Trained too long, some models turned overly harsh toward harmful requests, or tacked the same stock reassurance onto their answers to most red-team prompts (test prompts meant to coax out harmful answers). And needing fewer human labels makes it easier to ship a model that people have not tested closely. The authors say their aim is to make human oversight more efficient, not to drop it.
Standard preference training uses tens of thousands of human labels, and reading them tells you little about what the model is being trained to value.
Use written principles in both revision and preference training.
- 1 · specifyWrite a short list of plain-language principles that describe the behavior you want.
- 2 · reviseAsk a model to critique an initial response under a selected principle and produce a revision.
- 3 · tuneFine-tune the model on the revised responses.
- 4 · preferUse a model guided by the constitution to label response pairs, then train with those preferences.
A constitution makes intended values more legible, but the model must still interpret and generalize those words.
| Who | What they ask | What it works with |
|---|---|---|
| Safety team | “Which principle caused this refusal or revision?” | Critique traces and principle identifiers |
| Governance group | “Whose values shaped the constitution?” | Authoring and public-input process |
| Red team | “Can adversarial prompts bypass the trained behavior?” | Jailbreak success and over-refusal rates |
- Constitutional AI exposes behavioral principles as editable natural-language artifacts.
- It can generate critique-revision data and AI preference labels with fewer direct human labels.
- Public-input processes can be used to source some constitutional principles.
- Written principles can be incomplete, conflicting, or poorly interpreted.
- Someone still chooses the principles. The original paper picked its list in an ad hoc way, for research.
- It does not make a model immune to attacks, and the trained model may still not behave the way the document describes.
Sources used
This explainer is written in original language. The links below support its factual claims.
- paperConstitutional AI: Harmlessness from AI Feedback, Bai et al. · read 27 Sept 2026
- paperRLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback, Lee et al. · read 27 Sept 2026
- paperCollective Constitutional AI: Aligning a Language Model with Public Input, Huang et al. · read 27 Sept 2026
- docsClaude's Constitution, Anthropic · read 27 Sept 2026
- officialClaude's Constitution (2023 announcement of the original version), Anthropic · read 27 Sept 2026