LLM evaluation
LLM evaluation means testing an AI feature with structured tests, so you can check how accurate and reliable it is even though its answers vary.
A language model can give different replies to the same input, and that makes the usual ways of testing software fall short. So teams build evaluations, or “evals”: structured tests of how accurate and reliable an AI feature is.
This page is about LLM evaluation: testing features built on large language models.
Picture a bike shop chatbot. You want to try a new prompt, the written instructions the model gets before each chat. Is the new version better than the old one on real questions?
Start with real cases. Good test data can come from real chats with customers, from people who know the subject, and from old records. Mix everyday questions, edge cases and adversarial cases, which are questions written on purpose to trick the bot. An edge case is an unusual but possible input.
Say what passing means. A pass rule should be specific and measurable. Anthropic’s guide gives a loose goal as a warning: a model that should “classify sentiments well”, which roughly means be good at spotting happy and angry reviews. A tighter rule for the bike shop names the exact thing to look for: the answer states the return window from the policy page and does not invent one.
Choose a grader. A code check can compare the output with a known correct answer or look for a key phrase. Code checks are the fastest and most reliable kind of grading, but they miss the finer points.
When a known correct answer exists, fixed scoring formulas with names like ROUGE or BLEU can compare the output with it.
For fuzzier questions, a second model can grade the first. This is called an LLM judge. It is fast and handles harder calls, but test it before you rely on it. The judge needs a detailed, clear rubric: a marking sheet that spells out what earns a pass.
Google Cloud’s guide compares rubrics to unit tests. A unit test is a tiny automatic check a programmer writes for one small piece of code. A rubric does the same job for an answer: a short list of things that must be true. Some tools can even draft a custom checklist of yes-or-no tests for each prompt, which Google calls adaptive rubrics.
Can you trust a judge model? A 2023 research paper built a question set called MT-Bench and compared judge models with people. Strong LLM judges agreed with human preferences over 80% of the time. That is about as often as two humans agree with each other. Judge models showed biases, such as favouring an answer because of its position or its length.
OpenAI’s guidance says to use people’s own marks to check and adjust the automatic scores. Check a judge model agrees with people first; make it cheaper or faster later.
A worked example. Say the bike shop has twenty test questions, including “How do I return a helmet?” In the bike shop test, the code check is: does the answer contain the right return window? The judge follows a marking sheet for whether the tone is polite and the answer is clear. If it fails a case the old prompt passed, that is a regression: a step backwards. Regression tests can require a new version to beat the old one on the measures you care about.
Compare and trace. A trace is a record of all the in-between steps of one request, not just the final answer. When a case fails, you open its trace and see where it went wrong.
Offline and online. Offline evaluation is testing before launch. Online evaluation tracks quality on live traffic after launch. The two feed each other: issues found in live traffic can be added to the offline test set. Run evals on every change to catch a prompt that quietly makes things worse.
Language model features are hard to test the normal way.
Follow one change to a support bot through an evaluation run.
- 1 · collectGather test cases that mirror what real users send, including edge cases.
- 2 · runSend the same cases to the old version and the new version.
- 3 · gradeScore each answer with code checks, an LLM judge or human reviewers.
- 4 · comparePut the two sets of scores side by side and see which cases changed.
- 5 · traceOpen the traces of failed cases to see every intermediate step.
An evaluation is only as good as its grading: a vague rubric gives vague scores.
| Who | What they ask | What it works with |
|---|---|---|
| Support team | “Does the bot still pick the right label for refund questions?” | A code check that the label exactly matches the correct one |
| Product engineer | “Is a different model good enough for our summaries?” | Both models run on one dataset, compared side by side |
| Safety reviewer | “Are answers staying safe with real users after launch?” | Online evaluation of quality and safety on live traffic |
| Search team | “Do our answers meet the points on our marking sheet?” | Rubric scores from an LLM judge, checked against people's marks |
- It replaces a feeling that things work with checks on a fixed set of test cases.
- It catches regressions, where a new version does worse than the old one.
- It helps you pick the fastest, most reliable and most scalable way to grade.
- It feeds failures from live traffic back into the test set.
- Public benchmarks compare models in isolation, not your own feature.
- An LLM judge has biases, such as favouring an answer because of its position or length.
- Human grading is flexible and high quality, but slow and expensive.
- Online evaluation of live traffic usually has no reference answer to compare against.
Sources used
This explainer is written in original language. The links below support its factual claims.
- docsEvaluation best practices, OpenAI · read 28 Sept 2026
- docsDefine success criteria and build evaluations, Anthropic · read 28 Sept 2026
- docsGen AI evaluation service overview, Google Cloud · read 28 Sept 2026
- docsEvaluation concepts, LangChain · read 28 Sept 2026
- paperJudging LLM-as-a-Judge with MT-Bench and Chatbot Arena, arXiv (NeurIPS 2023 Datasets and Benchmarks Track) · read 28 Sept 2026