AI benchmarks
A benchmark is a fixed set of test tasks with a scoring rule, so different AI models can be compared on the same exam.
A benchmark is a test for AI models. It has a fixed set of tasks, such as questions or coding problems, and a fixed way to mark the answers. Because every model gets the same tasks and the same marking, their scores can be compared.
Benchmarks come in many shapes. MMLU asks questions across 57 subjects, from elementary maths to law. HumanEval asks a model to write small pieces of Python code, called functions, from a short note that says what each one should do. Then it checks whether the code works correctly. When it was released, the Codex model solved 28.8% of HumanEval problems, while GPT-3 solved none. Some suites, such as HELM, bundle many tests together. They measure more than right answers: they also check for bias, harmful or rude text, and how efficiently a model runs.
A score is only as good as the test behind it. A model learns from a huge pile of text called its training data. If the test questions leaked into that pile, the model may have memorised the answers instead of learning the skill. This problem is called contamination. One study showed that a 13-billion-parameter model could overfit a test this way. Its score jumped to about the same level as GPT-4. Strong scores can also hide gaps. The MMLU authors found that even the best models still fell short of experts on every subject.
Without a shared test, claims about models are hard to compare.
Follow one model through a benchmark run.
- 1 · fixPick a fixed set of tasks, such as the 57 subjects in MMLU.
- 2 · askGive the same prompts to each model.
- 3 · scoreMark each answer with the same rule, for example checking whether generated code works correctly.
- 4 · compareReport the scores side by side under the same conditions.
A score only means something next to the exact test and settings that produced it.
| Who | What they ask | What it works with |
|---|---|---|
| Model builder | “Did the new version get better at maths word problems?” | Scores on the same task before and after a change |
| App developer | “Which model writes working Python most often?” | Code benchmark results checked for functional correctness |
| Researcher | “Can others repeat my evaluation exactly?” | Public prompts, task configs and raw outputs |
- Gives many models the same tasks and the same scoring rule.
- Turns a vague question like "is it smart?" into a measurable score on named tasks.
- Public prompts and tools let others rerun the same evaluation.
- A high score on one test does not show a model is good at everything.
- If test questions leaked into training data, the score can be inflated.
- Accuracy alone can hide problems such as bias, toxicity or slowness.
Sources used
This explainer is written in original language. The links below support its factual claims.
- paperMeasuring Massive Multitask Language Understanding, Hendrycks and colleagues (arXiv) · read 28 Sept 2026
- paperHolistic Evaluation of Language Models, Liang and colleagues (arXiv) · read 28 Sept 2026
- paperEvaluating Large Language Models Trained on Code, Chen and colleagues (arXiv) · read 28 Sept 2026
- repoLanguage Model Evaluation Harness, EleutherAI · read 28 Sept 2026
- paperRethinking Benchmark and Contamination for Language Models with Rephrased Samples, Yang and colleagues (arXiv) · read 28 Sept 2026
- paperGLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding, Wang and colleagues (arXiv) · read 28 Sept 2026