Benchmark contamination
Benchmark contamination happens when test questions, or close copies of them, end up in a model's training data, so its score can look better than its real skill.
Benchmarks are test sets used to evaluate AI models. Benchmark contamination is when those test questions, or close copies of them, leak into the data a model learns from. Its score may then reflect what it has seen rather than real reasoning, like a student who saw the exam paper before the test. Contamination can quickly make a benchmark out of date.
A clear example comes from school maths. GSM1k was built to match GSM8k, a well-known set of grade-school maths word problems, but with new questions. On the new questions, some leading models scored up to 8% lower. But many of the most advanced models showed little sign of overfitting (doing well only on questions they had already met).
Researchers can spot leaks without seeing the training data. One test checks how likely the model finds the test questions in their original order, compared with shuffled orders. A model that saw the test tends to find the original order much more likely. Another test hides one wrong answer on a multiple-choice question, then asks the model to guess what was there. To stay fresh, LiveCodeBench, a coding test, keeps collecting new contest problems over time, and LiveBench adds and updates questions every month.
Public benchmarks can leak into the huge piles of text used for training.
Follow one test question from the internet into a score.
- 1 · publishMany benchmarks' test questions are public.
- 2 · leakData that closely resembles those questions ends up in a model's training data.
- 3 · memorizeThe model can partly memorize the questions, and even their order.
- 4 · compareTesting on fresh questions of the same style can reveal a drop in accuracy.
A high score on a leaked test measures memory as well as skill.
| Who | What they ask | What it works with |
|---|---|---|
| Benchmark builder | “How do I keep my questions out of future training sets?” | Fresh questions from recent sources, updated regularly |
| Model trainer | “Did any test questions sneak into our training data?” | Overlap checks between training data and benchmarks |
| Researcher | “Is this model reasoning or remembering?” | Scores on an old test versus a matched new test |
| Reader of a leaderboard | “Can I trust this headline number?” | Whether the benchmark is public and how old it is |
- Knowing about contamination helps explain why some scores overstate real ability.
- Matched fresh tests, such as GSM1k for grade-school maths, can measure the gap.
- Statistical tests can flag likely contamination without seeing the training data.
- Benchmarks built from recent material, such as LiveBench, limit the chance of leaks.
- String-matching filters alone do not remove reworded copies of test questions.
- Contamination can also arrive through synthetic data made by other models.
Sources used
This explainer is written in original language. The links below support its factual claims.
- paperLiveBench: A Challenging, Contamination-Limited LLM Benchmark, White and colleagues (arXiv) · read 28 Sept 2026
- paperA Careful Examination of Large Language Model Performance on Grade School Arithmetic, Zhang and colleagues (arXiv) · read 28 Sept 2026
- paperProving Test Set Contamination in Black Box Language Models, Oren and colleagues (arXiv) · read 28 Sept 2026
- paperRethinking Benchmark and Contamination for Language Models with Rephrased Samples, Yang and colleagues (arXiv) · read 28 Sept 2026
- paperLiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code, Jain and colleagues (arXiv) · read 28 Sept 2026
- paperInvestigating Data Contamination in Modern Benchmarks for Large Language Models, Deng and colleagues (arXiv) · read 28 Sept 2026