Agent evaluation
Agent evaluation tests an AI agent on set tasks, checking both its final result and the steps it took, over several tries.
An AI agent is a system where a language model decides its own steps and which tools to use. It works over many turns, calling tools and changing things as it goes. That makes an agent harder to test than a chatbot that gives one reply. One early mistake can carry forward and grow.
Agent evaluation gives the agent a fixed set of tasks, each with clear rules for success. Graders then score each attempt. Picture an agent that books flights. It may say the flight is booked, but the real test is whether a booking now exists in the test system’s database.
Because an agent can act differently each time, each task is run several times. The full record of a run is called a transcript or trace, and it shows every tool call and step. Reading these records shows whether the agent really failed or the grader was wrong.
For engineers: SWE-bench gives a model a code base (a project’s files) and a real GitHub issue (a reported bug or request) to fix. It has 2,294 such problems from 12 Python projects. τ-bench has a language model play the user while the agent uses tools and follows the rules it was given. It grades by comparing the database at the end of the conversation with the goal state. Its pass^k score asks a strict question: did the agent succeed on every one of k tries at the same task? AgentBench tests models as agents across 8 different environments. Google’s Agent Development Kit can score a run’s tool calls against an expected list.
Agents are harder to test than a single chat answer.
Follow one flight-booking task through an evaluation.
- 1 · taskWrite a task with a clear input and a rule for success.
- 2 · runRun the agent on the task several times, each from a clean start.
- 3 · gradeGraders check the final result and the steps the agent took to get there.
- 4 · scoreCombine the grades from every run into overall results.
- 5 · readRead the transcripts of failed runs to find the real cause.
Check what changed in the world, not only what the agent said.
| Who | What they ask | What it works with |
|---|---|---|
| Support team | “Did the agent really close the ticket, and in fewer than ten turns?” | A state check plus a limit on the number of turns |
| Coding team | “Does the fix pass the tests without breaking others?” | The project's own test suite |
| Chat agent team | “Does the agent succeed on every try, not just once?” | Several trials of the same simulated conversation |
| Agent developer | “Did the agent call the tools we expected?” | The list of tool calls compared with an expected list |
- It replaces guessing after each change with a repeatable check.
- Running several trials shows how reliable an agent is, not just whether it can succeed once.
- Checking the end state tests what really happened, not only what the agent said.
- Regression tests show when a change breaks a task the agent used to handle.
- Demanding one exact sequence of steps makes tests brittle.
- An agent can find a better answer that the test still marks as a failure.
- An LLM judge needs checking against human graders before you trust it.
- Leftover files or cached data between runs can make results wrong.
Sources used
This explainer is written in original language. The links below support its factual claims.
- officialDemystifying evals for AI agents, Anthropic · read 28 Sept 2026
- docsWhy evaluate agents, Google Agent Development Kit · read 28 Sept 2026
- paperτ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains, Yao et al. · read 28 Sept 2026
- paperSWE-bench: Can Language Models Resolve Real-World GitHub Issues?, Jimenez et al., ICLR 2024 · read 28 Sept 2026
- paperAgentBench: Evaluating LLMs as Agents, Liu et al., ICLR 2024 · read 28 Sept 2026
- officialBuilding effective agents, Anthropic · read 28 Sept 2026