Test set
A test set is a reserved group of labeled examples used for a final check of a model on data it did not train on.
A test set is holdout data, meaning the training procedure does not use it to fit model parameters. Its purpose is to measure the already-selected model on examples that stand in for future data. The test set holds both inputs and their correct answers, but the model only sees the inputs; a score such as accuracy then checks its predictions against the answers.
Keeping it apart matters because each score you look at can steer your next change. Look at test mistakes, tweak the model and test again a few times, and the model starts to fit that particular set. Tuning belongs to a validation set or to cross-validation, and the test set is kept out of that cycle.
It also needs enough examples for the score to be trusted, and none of them should be copies of training examples. Even a clean split only measures the data it holds, so a test set unlike real use cannot promise real-world results. When researchers built new ImageNet test sets by following the original collection process, a wide range of models lost 11% to 14% accuracy. The authors traced the drop to slightly harder images rather than to repeated use of the old test set.
Training scores cannot show whether a model will work on genuinely new examples.
Keep the final evaluation outside the development loop.
- 1 · splitBefore any training, set aside a slice of data that looks like what users will actually send the model.
- 2 · isolateLearn preprocessing and model parameters without fitting anything on the test examples.
- 3 · chooseUse training and validation results to settle the model and its hyperparameters.
- 4 · testGive the frozen model the test inputs only, then score its answers against the held-out labels.
If a test result changes what you build, that set has joined development and is no longer a clean final check.
| Who | What they ask | What it works with |
|---|---|---|
| Fraud team | “How well does the chosen model detect new cases?” | Final metrics on untouched transactions |
| Medical researcher | “Does the frozen classifier transfer to unseen patients?” | Patient-group holdout results |
| Forecasting team | “How accurate is the selected model on later dates?” | Time-ordered holdout results |
- An untouched set gives a fair estimate of how accurate the model is on the kind of data it was drawn from.
- A separate final evaluation keeps hyperparameter tuning from directly optimizing the reported test score.
- Group-aware splits can test whether a model generalizes to groups absent from training.
- A test set that differs from production data cannot prove production performance.
- Duplicate examples across training and test sets make the test unfair.
- Reusing the same test set for many adaptive choices can overfit the result to the holdout.
Sources used
This explainer is written in original language. The links below support its factual claims.
- docsDatasets: Dividing the original dataset, Google for Developers · read 27 Sept 2026
- docs3.1. Cross-validation: evaluating estimator performance, scikit-learn · read 27 Sept 2026
- docs12. Common pitfalls and recommended practices, scikit-learn · read 27 Sept 2026
- docsOverfitting, Google for Developers · read 27 Sept 2026
- paperGeneralization in Adaptive Data Analysis and Holdout Reuse, Dwork et al. · read 27 Sept 2026
- paperDo ImageNet Classifiers Generalize to ImageNet?, Recht et al., ICML 2019 · read 27 Sept 2026