Concepts

LLM-as-a-judge

3 min readintermediateUpdated 28 Sept 2026
1 · In one line

LLM-as-a-judge means asking a strong language model to grade answers to open-ended questions.

1 · What it is

LLM-as-a-judge means using one AI to test the answers of another. You give a language model (an AI that reads and writes text) an answer plus a rubric, which is a list of rules for a good answer. The judge returns a grade, such as a score from 1 to 5 or a simple correct or incorrect.

Imagine a homework helper bot that answered ten thousand questions. No teacher could mark them all by hand in a day, but a judge model can. The rubric can be the same set of scoring rules for every answer. It can also check things like whether the bot followed instructions, used the right format, or had the right tone and style. Researchers who tested the idea found it can come close to costly human ratings, at large scale, and it explains its scores.

The judge is not always fair. In tests it favoured the first answer it saw, longer answers and answers it wrote itself, and it may prefer machine-written text. Researchers check a judge by comparing its grades with human ratings.

2 · Why it exists

Grading open-ended answers is hard to do well at scale.

Humans are slowPeople are the most flexible, high-quality graders, but they cost a lot and take time.
Word matching missesOlder word-overlap scores like BLEU and ROUGE often disagree with people on creative tasks.
Answers are openChat assistants can do many things, and older tests measure what people prefer poorly.
3 · How it works

Follow one answer through a judge.

The judge follows a rubric and gives a score; researchers compare its scores with human ratings.
  1. 1 · rubricWrite a clear rubric that says what a good answer must contain.
  2. 2 · promptGive the judge model the rubric and the answer to grade.
  3. 3 · reasonAsk the judge to reason first, then give a fixed output such as correct or incorrect, or a score from 1 to 5.
  4. 4 · checkTest that the judge is reliable before using it at scale.

Before you rely on a judge, check its scores against human ratings.

4 · Where it's used
WhoWhat they askWhat it works with
Customer service team“Do our customer-service replies have the right tone”Replies scored on a 1-to-5 tone scale
Medical app team“Does any reply contain private health information”Each reply classified as yes or no
Writing tool team“Is our generated text good quality”Text scored with step-by-step reasoning and a form
5 · What it solves, and what it doesn't
solves
  • It grades far more answers than people can.
  • Strong judges agreed with human preferences over 80% of the time in one study.
  • It works for subjective qualities such as tone or empathy.
doesn't solve
  • Judges can favour the answer shown first, longer answers, or their own answers.
  • Judges have limited reasoning ability.
  • A judge may prefer text written by language models.
  • Building a reliable judge is still an open challenge.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. paperJudging LLM-as-a-Judge with MT-Bench and Chatbot Arena, arXiv · read 28 Sept 2026
  2. paperG-Eval: NLG Evaluation using GPT-4 with Better Human Alignment, arXiv · read 28 Sept 2026
  3. paperA Survey on LLM-as-a-Judge, arXiv · read 28 Sept 2026
  4. docsDefine success criteria and build evaluations, Anthropic · read 28 Sept 2026
  5. docsDefine your evaluation metrics, Google Cloud · read 28 Sept 2026