Concepts

Constitutional AI

4 min readintermediateUpdated 28 Sept 2026
1 · In one line

Constitutional AI trains and evaluates model behavior against an explicit written set of principles.

1 · What it is

Constitutional AI makes a set of behavioral principles an explicit part of post-training. In its original form, a model first generates a response, critiques that response under a sampled principle, and revises it. Those revisions become supervised fine-tuning data.

In a second phase, the fine-tuned model writes two answers to the same prompt. An AI judge picks the better one, guided by a principle from the constitution. A preference model learns from many of those picks, mixed with human labels for helpfulness, and reinforcement learning then uses it to score the assistant. The paper calls this reinforcement learning from AI feedback, or RLAIF. A later study by a different team tried RLAIF on three jobs: summarizing text, giving helpful chat answers, and giving harmless ones. On all three, it did about as well as training on human labels (RLHF).

The constitution makes intended behavior easier to inspect and revise, but it does not mechanically determine outputs. The model has to interpret the words and carry them over to cases the authors did not foresee, and even Anthropic expects parts of Claude’s behavior to fall short of its own constitution. The original authors also flag risks. Trained too long, some models turned overly harsh toward harmful requests, or tacked the same stock reassurance onto their answers to most red-team prompts (test prompts meant to coax out harmful answers). And needing fewer human labels makes it easier to ship a model that people have not tested closely. The authors say their aim is to make human oversight more efficient, not to drop it.

2 · Why it exists

Standard preference training uses tens of thousands of human labels, and reading them tells you little about what the model is being trained to value.

Implicit rulesNobody can read every label and work out what they add up to.
Label scaleAs models get more capable, training may need methods where people do not check every part of their behavior. AI feedback may be a more efficient way to supervise.
GovernanceDevelopers alone may not represent the communities affected by model behavior.
3 · How it works

Use written principles in both revision and preference training.

Constitutional AI uses principles for revision and preference labelsA constitution guides a highlighted critique and revision path for supervised tuning, and separately guides an AI judge that creates preference data for reinforcement learning. Initial answerfrom user promptmay violate a rule Constitutionprinciple: explain harmwithout enabling it KEY STEPCritique + reviseidentify conflictrewrite the answerprinciple stays visible Revised answerSFT exampletrain on revision AI preference judgesample A vs sample Bconstitution → chooseRLAIF preference dataPreference model + RLhybrid human/AI labelsThe same written principles guide two different data-generation paths.
The constitution is an inspectable input to two training stages: self-revision and AI-generated preferences.
  1. 1 · specifyWrite a short list of plain-language principles that describe the behavior you want.
  2. 2 · reviseAsk a model to critique an initial response under a selected principle and produce a revision.
  3. 3 · tuneFine-tune the model on the revised responses.
  4. 4 · preferUse a model guided by the constitution to label response pairs, then train with those preferences.

A constitution makes intended values more legible, but the model must still interpret and generalize those words.

4 · Where it's used
WhoWhat they askWhat it works with
Safety team“Which principle caused this refusal or revision?”Critique traces and principle identifiers
Governance group“Whose values shaped the constitution?”Authoring and public-input process
Red team“Can adversarial prompts bypass the trained behavior?”Jailbreak success and over-refusal rates
5 · What it solves, and what it doesn't
solves
  • Constitutional AI exposes behavioral principles as editable natural-language artifacts.
  • It can generate critique-revision data and AI preference labels with fewer direct human labels.
  • Public-input processes can be used to source some constitutional principles.
doesn't solve
  • Written principles can be incomplete, conflicting, or poorly interpreted.
  • Someone still chooses the principles. The original paper picked its list in an ad hoc way, for research.
  • It does not make a model immune to attacks, and the trained model may still not behave the way the document describes.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. paperConstitutional AI: Harmlessness from AI Feedback, Bai et al. · read 27 Sept 2026
  2. paperRLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback, Lee et al. · read 27 Sept 2026
  3. paperCollective Constitutional AI: Aligning a Language Model with Public Input, Huang et al. · read 27 Sept 2026
  4. docsClaude's Constitution, Anthropic · read 27 Sept 2026
  5. officialClaude's Constitution (2023 announcement of the original version), Anthropic · read 27 Sept 2026