Self-reflection
Self-reflection is when an AI model checks its own first answer, writes feedback on it, and then tries again using that feedback.
Language models do not always get things right on the first try. Self-reflection gives the model a second look, like a writer rereading a draft. The model writes an answer. Then it gives itself feedback on that answer. Then it uses the feedback to fix the answer, and it can repeat these steps. Across seven test tasks in the Self-Refine study, the loop scored about 20 percentage points higher on average than the same model’s one-try answers.
The Reflexion method stores the model’s written reflections in memory for later attempts. The model’s weights (the numbers it learned in training) do not change. Picture a coding agent: it can read feedback on a failed try and note what to change. Reflexion reached 91 percent pass@1 on the HumanEval coding benchmark, against 80 percent for GPT-4.
Feedback from outside the model matters. A method called CRITIC lets a model check and fix its drafts by using outside tools. A later study looked at reasoning tasks. Without outside feedback, models struggled to fix their own answers, and sometimes got worse. Anthropic describes a related pattern: one model call writes, another checks the result, and they loop. It helps most when there are clear rules for a good answer.
A model's first answer is not always its best one.
Follow one coding task through a reflection loop.
- 1 · draftThe model writes a first answer to the task.
- 2 · checkFeedback arrives, from the model itself or from an outside tool such as a code interpreter.
- 3 · reflectThe model writes plain-language notes about what went wrong.
- 4 · retryThose notes are kept and given to the model on its next attempt.
Self-reflection improves answers by adding notes to the prompt, not by retraining.
| Who | What they ask | What it works with |
|---|---|---|
| Coding agent | “Why did my function fail two of the tests?” | Its own code plus the result of running it |
| Translator | “Does this translation keep the tone of the original?” | A draft translation and a critique of it |
| Research assistant | “Do I need to search again before I answer?” | Its own summary and what is still missing |
- It can improve answers without any extra training.
- One model can write, critique and rewrite its own work.
- With tests or search results as feedback, it can catch real errors.
- Without outside feedback, models often fail to fix their own reasoning.
- Sometimes the answer gets worse after the model "corrects" it.
- Each round needs another model call.
Sources used
This explainer is written in original language. The links below support its factual claims.
- paperReflexion: Language Agents with Verbal Reinforcement Learning, Shinn et al. · read 28 Sept 2026
- paperSelf-Refine: Iterative Refinement with Self-Feedback, Madaan et al. · read 28 Sept 2026
- paperCRITIC: Large Language Models Can Self-Correct with Tool-Interactive Critiquing, Gou et al. · read 28 Sept 2026
- paperLarge Language Models Cannot Self-Correct Reasoning Yet, Huang et al. · read 28 Sept 2026
- officialBuilding effective agents, Anthropic · read 28 Sept 2026