BenchmarksConcepts

AI benchmarks

3 min readbeginnerUpdated 28 Sept 2026
1 · In one line

A benchmark is a fixed set of test tasks with a scoring rule, so different AI models can be compared on the same exam.

1 · What it is

A benchmark is a test for AI models. It has a fixed set of tasks, such as questions or coding problems, and a fixed way to mark the answers. Because every model gets the same tasks and the same marking, their scores can be compared.

Benchmarks come in many shapes. MMLU asks questions across 57 subjects, from elementary maths to law. HumanEval asks a model to write small pieces of Python code, called functions, from a short note that says what each one should do. Then it checks whether the code works correctly. When it was released, the Codex model solved 28.8% of HumanEval problems, while GPT-3 solved none. Some suites, such as HELM, bundle many tests together. They measure more than right answers: they also check for bias, harmful or rude text, and how efficiently a model runs.

A score is only as good as the test behind it. A model learns from a huge pile of text called its training data. If the test questions leaked into that pile, the model may have memorised the answers instead of learning the skill. This problem is called contamination. One study showed that a 13-billion-parameter model could overfit a test this way. Its score jumped to about the same level as GPT-4. Strong scores can also hide gaps. The MMLU authors found that even the best models still fell short of experts on every subject.

2 · Why it exists

Without a shared test, claims about models are hard to compare.

Different testsBefore HELM, its authors found that some well-known models did not share even one test scenario.
Narrow skillsGLUE was built to check language understanding across many existing tasks, not just one dataset.
Hidden weak spotsThe MMLU authors say their test can help find important shortcomings across many subjects.
3 · How it works

Follow one model through a benchmark run.

One benchmark run: fixed questions go to a model, a fixed scoring rule marks the answers, and the result is a score comparable with other models ONE BENCHMARK RUN Fixed test set Model Score Scoring rule same questions for every model e.g. 57 subjects reads prompt writes answer or code share of tasks marked correct e.g. 28.8% solved match answer key or check the code same rule for all Compare side by side model A vs model B, same test and settings
The scoring rule stays the same for every model, which is what makes scores comparable.
  1. 1 · fixPick a fixed set of tasks, such as the 57 subjects in MMLU.
  2. 2 · askGive the same prompts to each model.
  3. 3 · scoreMark each answer with the same rule, for example checking whether generated code works correctly.
  4. 4 · compareReport the scores side by side under the same conditions.

A score only means something next to the exact test and settings that produced it.

4 · Where it's used
WhoWhat they askWhat it works with
Model builder“Did the new version get better at maths word problems?”Scores on the same task before and after a change
App developer“Which model writes working Python most often?”Code benchmark results checked for functional correctness
Researcher“Can others repeat my evaluation exactly?”Public prompts, task configs and raw outputs
5 · What it solves, and what it doesn't
solves
  • Gives many models the same tasks and the same scoring rule.
  • Turns a vague question like "is it smart?" into a measurable score on named tasks.
  • Public prompts and tools let others rerun the same evaluation.
doesn't solve
  • A high score on one test does not show a model is good at everything.
  • If test questions leaked into training data, the score can be inflated.
  • Accuracy alone can hide problems such as bias, toxicity or slowness.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. paperMeasuring Massive Multitask Language Understanding, Hendrycks and colleagues (arXiv) · read 28 Sept 2026
  2. paperHolistic Evaluation of Language Models, Liang and colleagues (arXiv) · read 28 Sept 2026
  3. paperEvaluating Large Language Models Trained on Code, Chen and colleagues (arXiv) · read 28 Sept 2026
  4. repoLanguage Model Evaluation Harness, EleutherAI · read 28 Sept 2026
  5. paperRethinking Benchmark and Contamination for Language Models with Rephrased Samples, Yang and colleagues (arXiv) · read 28 Sept 2026
  6. paperGLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding, Wang and colleagues (arXiv) · read 28 Sept 2026