Building with AI

Agent evaluation

4 min readintermediateUpdated 28 Sept 2026
1 · In one line

Agent evaluation tests an AI agent on set tasks, checking both its final result and the steps it took, over several tries.

1 · What it is

An AI agent is a system where a language model decides its own steps and which tools to use. It works over many turns, calling tools and changing things as it goes. That makes an agent harder to test than a chatbot that gives one reply. One early mistake can carry forward and grow.

Agent evaluation gives the agent a fixed set of tasks, each with clear rules for success. Graders then score each attempt. Picture an agent that books flights. It may say the flight is booked, but the real test is whether a booking now exists in the test system’s database.

Because an agent can act differently each time, each task is run several times. The full record of a run is called a transcript or trace, and it shows every tool call and step. Reading these records shows whether the agent really failed or the grader was wrong.

For engineers: SWE-bench gives a model a code base (a project’s files) and a real GitHub issue (a reported bug or request) to fix. It has 2,294 such problems from 12 Python projects. τ-bench has a language model play the user while the agent uses tools and follows the rules it was given. It grades by comparing the database at the end of the conversation with the goal state. Its pass^k score asks a strict question: did the agent succeed on every one of k tries at the same task? AgentBench tests models as agents across 8 different environments. Google’s Agent Development Kit can score a run’s tool calls against an expected list.

2 · Why it exists

Agents are harder to test than a single chat answer.

Many stepsAn agent calls tools over many turns, so one early mistake can carry forward and grow.
Different each runA task that passes on one run can fail on the next.
Words are not resultsWhat the agent says at the end is not the test; what actually changed is.
3 · How it works

Follow one flight-booking task through an evaluation.

The grading step checks what really changed, not only what the agent said.
  1. 1 · taskWrite a task with a clear input and a rule for success.
  2. 2 · runRun the agent on the task several times, each from a clean start.
  3. 3 · gradeGraders check the final result and the steps the agent took to get there.
  4. 4 · scoreCombine the grades from every run into overall results.
  5. 5 · readRead the transcripts of failed runs to find the real cause.

Check what changed in the world, not only what the agent said.

4 · Where it's used
WhoWhat they askWhat it works with
Support team“Did the agent really close the ticket, and in fewer than ten turns?”A state check plus a limit on the number of turns
Coding team“Does the fix pass the tests without breaking others?”The project's own test suite
Chat agent team“Does the agent succeed on every try, not just once?”Several trials of the same simulated conversation
Agent developer“Did the agent call the tools we expected?”The list of tool calls compared with an expected list
5 · What it solves, and what it doesn't
solves
  • It replaces guessing after each change with a repeatable check.
  • Running several trials shows how reliable an agent is, not just whether it can succeed once.
  • Checking the end state tests what really happened, not only what the agent said.
  • Regression tests show when a change breaks a task the agent used to handle.
doesn't solve
  • Demanding one exact sequence of steps makes tests brittle.
  • An agent can find a better answer that the test still marks as a failure.
  • An LLM judge needs checking against human graders before you trust it.
  • Leftover files or cached data between runs can make results wrong.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. officialDemystifying evals for AI agents, Anthropic · read 28 Sept 2026
  2. docsWhy evaluate agents, Google Agent Development Kit · read 28 Sept 2026
  3. paperτ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains, Yao et al. · read 28 Sept 2026
  4. paperSWE-bench: Can Language Models Resolve Real-World GitHub Issues?, Jimenez et al., ICLR 2024 · read 28 Sept 2026
  5. paperAgentBench: Evaluating LLMs as Agents, Liu et al., ICLR 2024 · read 28 Sept 2026
  6. officialBuilding effective agents, Anthropic · read 28 Sept 2026