Data leakage
Data leakage is when a model learns from information it will not have when making real predictions, so its test score looks better than it really is.
Data leakage happens when a model is built using information it will not have when it makes real predictions. The model can then take a shortcut: it uses the leaked detail instead of learning the real pattern. Testing rewards the shortcut, so the score comes out too optimistic, and the model does worse on genuinely new data once it is in use.
The first kind, called label leakage, is a feature that stands in for the answer. Google’s crash course describes a model meant to spot cancer in new hospital patients. It did brilliantly in testing and badly in real use. The reason was the hospital name: patients were sent to a specialist hospital only after diagnosis, so at prediction time most had no hospital yet. A retail version is predicting a day’s revenue from that day’s customer count, which nobody knows until the day’s sales are done.
The second kind lets test data shape training. If a step such as scaling, filling in missing values or choosing features is learned from the whole dataset before splitting, the test set has already influenced the model; the feature scaling entry shows one case. In scikit-learn’s demonstration the labels were random, so about 50 per cent accuracy was expected. Choosing features before splitting pushed accuracy far above chance; splitting first brought it back down. Duplicate rows across the split and shuffled time-ordered data, which lets a model train on the future, leak the same way. For language models the version to watch is benchmark contamination, where test questions turn up in web-crawled training data.
The damage is real. A 2023 survey found leakage in at least 294 papers across 17 fields. In civil war prediction, the four papers claiming complex models beat logistic regression all had leaks. Once fixed, the complex models were no better. A suspiciously near-perfect score is a reason to look for a leak.
A leaked score tells you the wrong thing about your model.
Follow one leaky column from the table to the test score.
- 1 · splitSplit the data into training and test sets first, before any preprocessing.
- 2 · timeFor time-ordered data, train on earlier records and test on later ones instead of shuffling.
- 3 · checkKeep only features you will actually have when the model is used.
- 4 · fitLearn every preprocessing step from the training set only, ideally inside a pipeline.
- 5 · dedupeRemove test examples that duplicate training examples.
Ask of every feature: would I know this value at the moment of prediction?
| Who | What they ask | What it works with |
|---|---|---|
| Hospital data team | “Does any column get filled in only after the diagnosis is made?” | When each feature is recorded compared with the outcome |
| Retail forecasting team | “Will we know today's customer count before the day's sales are done?” | Which features exist at the time of the forecast |
| Research lab | “Did we fill in missing values before or after splitting the data?” | The order of preprocessing and splitting in the code |
| Language model team | “Did benchmark questions end up in our web-crawled training text?” | Overlap between training data and test sets |
- Splitting first keeps test data out of every choice made about the model.
- A pipeline learns preprocessing from training data only, including inside cross-validation.
- Time-based splits stop a model from training on the future and testing on the past.
- Checking when each feature is recorded catches columns that give the answer away.
- Some leaks stay hidden; a feature that stands in for the label can be very hard to detect.
- A clean split does not prove the test set matches the situation you care about.
- For web-scale language models, there were no settled methods for detecting test contamination.
- Fixing a leak does not make a model better; it can erase an advantage that was never real.
Sources used
This explainer is written in original language. The links below support its factual claims.
- docsCommon pitfalls and recommended practices (scikit-learn user guide), scikit-learn · read 27 Sept 2026
- docsTimeSeriesSplit, scikit-learn · read 27 Sept 2026
- docsMachine Learning Glossary, Google for Developers · read 27 Sept 2026
- docsProduction ML systems: Monitoring pipelines (Machine Learning Crash Course), Google for Developers · read 27 Sept 2026
- docsDatasets: Dividing the original dataset (Machine Learning Crash Course), Google for Developers · read 27 Sept 2026
- paperLeakage and the reproducibility crisis in machine-learning-based science, Patterns (Cell Press, 2023; Kapoor and Narayanan), via PubMed Central · read 27 Sept 2026
- paperLanguage Models are Few-Shot Learners, arXiv (OpenAI, Brown et al.) · read 27 Sept 2026