LLM-as-a-judge
LLM-as-a-judge means asking a strong language model to grade answers to open-ended questions.
LLM-as-a-judge means using one AI to test the answers of another. You give a language model (an AI that reads and writes text) an answer plus a rubric, which is a list of rules for a good answer. The judge returns a grade, such as a score from 1 to 5 or a simple correct or incorrect.
Imagine a homework helper bot that answered ten thousand questions. No teacher could mark them all by hand in a day, but a judge model can. The rubric can be the same set of scoring rules for every answer. It can also check things like whether the bot followed instructions, used the right format, or had the right tone and style. Researchers who tested the idea found it can come close to costly human ratings, at large scale, and it explains its scores.
The judge is not always fair. In tests it favoured the first answer it saw, longer answers and answers it wrote itself, and it may prefer machine-written text. Researchers check a judge by comparing its grades with human ratings.
Grading open-ended answers is hard to do well at scale.
Follow one answer through a judge.
- 1 · rubricWrite a clear rubric that says what a good answer must contain.
- 2 · promptGive the judge model the rubric and the answer to grade.
- 3 · reasonAsk the judge to reason first, then give a fixed output such as correct or incorrect, or a score from 1 to 5.
- 4 · checkTest that the judge is reliable before using it at scale.
Before you rely on a judge, check its scores against human ratings.
| Who | What they ask | What it works with |
|---|---|---|
| Customer service team | “Do our customer-service replies have the right tone” | Replies scored on a 1-to-5 tone scale |
| Medical app team | “Does any reply contain private health information” | Each reply classified as yes or no |
| Writing tool team | “Is our generated text good quality” | Text scored with step-by-step reasoning and a form |
- It grades far more answers than people can.
- Strong judges agreed with human preferences over 80% of the time in one study.
- It works for subjective qualities such as tone or empathy.
- Judges can favour the answer shown first, longer answers, or their own answers.
- Judges have limited reasoning ability.
- A judge may prefer text written by language models.
- Building a reliable judge is still an open challenge.
Sources used
This explainer is written in original language. The links below support its factual claims.
- paperJudging LLM-as-a-Judge with MT-Bench and Chatbot Arena, arXiv · read 28 Sept 2026
- paperG-Eval: NLG Evaluation using GPT-4 with Better Human Alignment, arXiv · read 28 Sept 2026
- paperA Survey on LLM-as-a-Judge, arXiv · read 28 Sept 2026
- docsDefine success criteria and build evaluations, Anthropic · read 28 Sept 2026
- docsDefine your evaluation metrics, Google Cloud · read 28 Sept 2026