Evaluations (evals)
An eval checks an AI system's work. You feed it an input and use a set of rules to score what comes back.
An eval, short for evaluation, checks an AI system’s work. You feed it an input, score the reply with a set of rules, and note whether it passed. Google compares its eval checklists to unit tests. A unit test is a small automatic check that programmers write to make sure their code still works.
Why guessing is not enough
A model is the AI program that writes the answers. OpenAI points out that without evals, it is slow and hard to learn what a new version does to your own app. Anthropic adds that teams without evals cannot tell real regressions from noise. A regression is something that used to work and now fails.
The three parts of an eval
Every eval needs a clear idea of success, a set of test cases, and a grader.
Success. Anthropic’s first tip is to spell out exactly what a good result looks like.
Test cases. Anthropic says your test cases should look like the mix of jobs the system will really face. It also reminds you not to forget edge cases, the odd inputs that rarely show up. Its examples include irrelevant or missing input and very long input.
A grader. This is the part that decides whether each output passes. In OpenAI’s evals, the graders are what decide whether an answer counts as right. There are several kinds:
- Exact match checks whether the output is the same as a correct answer you wrote down ahead of time.
- A rubric is the list of criteria for rating a response, such as “stays on topic” or “is polite”.
- A person can judge the answer. Anthropic rates people as the most adaptable, highest-quality graders, though they are slow and costly.
- Another model can act as a judge. Anthropic says this is quick, adaptable, easy to scale and good for tricky calls, but you should check that it judges well before you lean on it for lots of tests.
Anthropic’s advice is to write questions so a computer can mark them. Examples are matching exact text, multiple choice, marking by a small program, or marking by a model. Its view: lots of questions marked by a machine, even a bit less exactly, beat a few marked by hand.
A worked example
Imagine a school helpdesk bot that sorts student messages into four moods: negative, positive, mixed or neutral. You write 200 real-looking messages and label each one, including sarcastic ones like “Great, another surprise quiz.” Exact match grading gives the bot a score.
Now you rewrite its instructions (the prompt), or switch to a newer model. You run the same 200 messages again. If the sarcasm cases improve but the mixed messages now fail, you have found a regression before any student sees it. Say v1 scored 181 of 200 and v2 scores 188: better overall, yet check what changed.
Add odd messages too: a blank one, a pasted lunch menu, a page-long rant. Score the share handled with no errors, say 48 of 50, or 96%.
Running evals again and again
OpenAI describes a loop: run the eval, study what went wrong, change your prompt, and run it again.
Anthropic separates two kinds of evals for this loop. Capability evals ask what a system can do well. They aim at tasks the system still finds hard, so they start with a low pass rate. That gives the team something to climb. Regression evals check that old skills have not broken, so they should pass almost every time. Once a capability eval passes most of the time, it can graduate and join the set of regression evals.
Where evals fall short
In one study, strong model judges agreed with people’s choices over 80% of the time. That is about as often as people agree with each other. But the same study found the judges had built-in leanings. A judge can favour the first answer, the longer one, or its own. Anthropic separately advises testing a model judge for reliability before scaling it up.
Without a fixed test, nobody can tell whether a change helped or hurt.
Follow one change to a support bot through an eval.
- 1 · defineDescribe the task and what counts as a correct answer.
- 2 · collectGather test inputs that look like the real work, including tricky edge cases.
- 3 · runSend every test input through the system and save each output.
- 4 · gradeA grader checks each output by exact match, a rubric, a person or another model.
- 5 · compareCompare the new score with the last one, then improve and run again.
If a model does the grading, test the grader for reliability first, then scale it up.
| Who | What they ask | What it works with |
|---|---|---|
| App developer | “Does the new model version still answer our customers correctly?” | The same test set, run on the old and new model |
| Prompt writer | “Did my rewritten instructions fix the sarcasm cases without breaking others?” | Labelled examples, graded by exact match |
| Agent team | “Can the agent still finish every task it used to finish?” | A regression suite of past tasks |
| Product team | “Are long answers polite and on topic?” | A rubric applied by a model judge, checked against human ratings |
- Scores each answer against set rules to see if it worked.
- Helps teams tell real regressions from noise.
- Helps you understand how a new model version might affect your use case.
- Forces you to spell out what a good result looks like.
- Some test cases are so ambiguous that even people would disagree on the grade.
- A model judge has its own biases, such as favouring longer answers.
- Human grading is high quality but slow and expensive.
Sources used
This explainer is written in original language. The links below support its factual claims.
- docsDefine success criteria and build evaluations, Anthropic · read 28 Sept 2026
- officialDemystifying evals for AI agents, Anthropic · read 28 Sept 2026
- docsWorking with evals, OpenAI · read 28 Sept 2026
- docsGen AI evaluation service overview, Google Cloud · read 28 Sept 2026
- repoopenai/evals, OpenAI · read 28 Sept 2026
- paperJudging LLM-as-a-Judge with MT-Bench and Chatbot Arena, arXiv · read 28 Sept 2026