EvalsConcepts

Evaluations (evals)

5 min readbeginnerUpdated 28 Sept 2026
1 · In one line

An eval checks an AI system's work. You feed it an input and use a set of rules to score what comes back.

1 · What it is

An eval, short for evaluation, checks an AI system’s work. You feed it an input, score the reply with a set of rules, and note whether it passed. Google compares its eval checklists to unit tests. A unit test is a small automatic check that programmers write to make sure their code still works.

Why guessing is not enough

A model is the AI program that writes the answers. OpenAI points out that without evals, it is slow and hard to learn what a new version does to your own app. Anthropic adds that teams without evals cannot tell real regressions from noise. A regression is something that used to work and now fails.

The three parts of an eval

Every eval needs a clear idea of success, a set of test cases, and a grader.

Success. Anthropic’s first tip is to spell out exactly what a good result looks like.

Test cases. Anthropic says your test cases should look like the mix of jobs the system will really face. It also reminds you not to forget edge cases, the odd inputs that rarely show up. Its examples include irrelevant or missing input and very long input.

A grader. This is the part that decides whether each output passes. In OpenAI’s evals, the graders are what decide whether an answer counts as right. There are several kinds:

  • Exact match checks whether the output is the same as a correct answer you wrote down ahead of time.
  • A rubric is the list of criteria for rating a response, such as “stays on topic” or “is polite”.
  • A person can judge the answer. Anthropic rates people as the most adaptable, highest-quality graders, though they are slow and costly.
  • Another model can act as a judge. Anthropic says this is quick, adaptable, easy to scale and good for tricky calls, but you should check that it judges well before you lean on it for lots of tests.

Anthropic’s advice is to write questions so a computer can mark them. Examples are matching exact text, multiple choice, marking by a small program, or marking by a model. Its view: lots of questions marked by a machine, even a bit less exactly, beat a few marked by hand.

A worked example

Imagine a school helpdesk bot that sorts student messages into four moods: negative, positive, mixed or neutral. You write 200 real-looking messages and label each one, including sarcastic ones like “Great, another surprise quiz.” Exact match grading gives the bot a score.

Now you rewrite its instructions (the prompt), or switch to a newer model. You run the same 200 messages again. If the sarcasm cases improve but the mixed messages now fail, you have found a regression before any student sees it. Say v1 scored 181 of 200 and v2 scores 188: better overall, yet check what changed.

Add odd messages too: a blank one, a pasted lunch menu, a page-long rant. Score the share handled with no errors, say 48 of 50, or 96%.

Running evals again and again

OpenAI describes a loop: run the eval, study what went wrong, change your prompt, and run it again.

Anthropic separates two kinds of evals for this loop. Capability evals ask what a system can do well. They aim at tasks the system still finds hard, so they start with a low pass rate. That gives the team something to climb. Regression evals check that old skills have not broken, so they should pass almost every time. Once a capability eval passes most of the time, it can graduate and join the set of regression evals.

Where evals fall short

In one study, strong model judges agreed with people’s choices over 80% of the time. That is about as often as people agree with each other. But the same study found the judges had built-in leanings. A judge can favour the first answer, the longer one, or its own. Anthropic separately advises testing a model judge for reliability before scaling it up.

2 · Why it exists

Without a fixed test, nobody can tell whether a change helped or hurt.

Upgrades are guessworkWithout evals, it is slow and hard to see how a new model version would affect your own use case.
Fixes break thingsTeams without evals cannot tell real regressions, meaning things that used to work and now fail, from random noise.
Errors pile upAn agent that uses tools over many steps can make one mistake that spreads into later steps.
3 · How it works

Follow one change to a support bot through an eval.

The dataset and grader stay fixed, so the score only moves when the system changes.
  1. 1 · defineDescribe the task and what counts as a correct answer.
  2. 2 · collectGather test inputs that look like the real work, including tricky edge cases.
  3. 3 · runSend every test input through the system and save each output.
  4. 4 · gradeA grader checks each output by exact match, a rubric, a person or another model.
  5. 5 · compareCompare the new score with the last one, then improve and run again.

If a model does the grading, test the grader for reliability first, then scale it up.

4 · Where it's used
WhoWhat they askWhat it works with
App developer“Does the new model version still answer our customers correctly?”The same test set, run on the old and new model
Prompt writer“Did my rewritten instructions fix the sarcasm cases without breaking others?”Labelled examples, graded by exact match
Agent team“Can the agent still finish every task it used to finish?”A regression suite of past tasks
Product team“Are long answers polite and on topic?”A rubric applied by a model judge, checked against human ratings
5 · What it solves, and what it doesn't
solves
  • Scores each answer against set rules to see if it worked.
  • Helps teams tell real regressions from noise.
  • Helps you understand how a new model version might affect your use case.
  • Forces you to spell out what a good result looks like.
doesn't solve
  • Some test cases are so ambiguous that even people would disagree on the grade.
  • A model judge has its own biases, such as favouring longer answers.
  • Human grading is high quality but slow and expensive.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. docsDefine success criteria and build evaluations, Anthropic · read 28 Sept 2026
  2. officialDemystifying evals for AI agents, Anthropic · read 28 Sept 2026
  3. docsWorking with evals, OpenAI · read 28 Sept 2026
  4. docsGen AI evaluation service overview, Google Cloud · read 28 Sept 2026
  5. repoopenai/evals, OpenAI · read 28 Sept 2026
  6. paperJudging LLM-as-a-Judge with MT-Bench and Chatbot Arena, arXiv · read 28 Sept 2026