Validation set
A validation set is held-out data used during development to compare model choices without training on the same examples.
The validation set is separate from the training set, but its results are part of model development. Developers compare candidates on it, change hyperparameters or features and run the cycle again. That feedback makes validation useful for selection and unsuitable as a final unbiased report.
The simplest setup is one fixed split, where the model learns from the training part and is then scored on the validation part. With k-fold cross-validation, each fold becomes validation once while the other folds train the model. Rotating like this wastes less data, which helps most when there are only a few examples to go around.
The split must respect how data is related. Several readings from one patient belong together, so they go on the same side of the split, and data collected over time needs a split that respects time order. Once development ends, the selected model moves to a separate test set.
Developers need feedback for choosing a model without turning the final test into another tuning tool.
Put validation inside the development loop and testing after it.
- 1 · fitTrain each candidate only on the training partition.
- 2 · scoreMeasure each candidate on validation examples it did not fit.
- 3 · compareUse the validation score to choose hyperparameters, features or a stopping point.
- 4 · repeatIterate until the development rule is satisfied, then freeze the choice for final testing.
Validation is allowed to influence the model; that is exactly why a separate test set still matters.
| Who | What they ask | What it works with |
|---|---|---|
| Model trainer | “Which learning rate should this run use?” | Validation loss for each candidate rate |
| Vision engineer | “Which augmentation recipe transfers best?” | Validation score for each recipe |
| NLP researcher | “When should training stop?” | Validation loss across training steps |
- Validation gives a score for choosing hyperparameters without fitting those choices on the training examples.
- Diverging training and validation loss curves can reveal overfitting.
- Cross-validation rotates which training fold acts as validation when a fixed validation set would waste scarce data.
- Repeated validation use can make model choices fit that validation set.
- Validation does not replace an untouched test set for final evaluation.
- Random folds are inappropriate when time order or dependent groups would leak information across the split.
Sources used
This explainer is written in original language. The links below support its factual claims.
- docsDatasets: Dividing the original dataset, Google for Developers · read 27 Sept 2026
- docs3.1. Cross-validation: evaluating estimator performance, scikit-learn · read 27 Sept 2026
- docs3.5. Validation curves: plotting scores to evaluate models, scikit-learn · read 27 Sept 2026
- docsOverfitting, Google for Developers · read 27 Sept 2026
- paperGeneralization in Adaptive Data Analysis and Holdout Reuse, Dwork et al. · read 27 Sept 2026
- paperOn Over-fitting in Model Selection and Subsequent Selection Bias in Performance Evaluation, Cawley and Talbot, Journal of Machine Learning Research · read 27 Sept 2026