Missing data
Missing data are empty cells in a dataset. You can drop them, fill them with reasoned guesses or use a model that accepts gaps, and why they are missing decides which is safe.
Missing data are cells in a dataset that hold no value. Real datasets often have them, stored as blanks, NaN (short for “not a number”) or another placeholder. In pandas, the marker for a missing value, called NA, depends on the column’s data type. A made-up code such as minus one for “no measurement” is worse, because the model treats it as a real number.
How the gaps arose matters as much as how many there are. Rubin sorted missing data into three categories in 1976. Missing completely at random, or MCAR, means every value had the same chance of going missing, so the gaps are unrelated to the data. A weighing scale whose battery ran out is the classic example. Missing at random, or MAR, means the chance differs only between groups you can see in the data. A scale that fails more often on soft floors fits, as long as the floor type was recorded. Missing not at random, or MNAR, means the chance depends on reasons you cannot see, such as people with weaker opinions answering a poll less often. Most modern methods assume at least MAR.
Deleting rows is the bluntest option, and under MCAR it still gives unbiased means and regression weights. In pandas, dropna removes rows or columns that contain gaps. Dropping a whole column makes sense when one feature is mostly empty and adds little. Google’s crash course suggests building one dataset each way and keeping whichever trains the better model.
Imputation means filling the gaps with reasoned guesses. Simple imputation fills each blank with a column’s mean, median or most frequent value, or with a fixed constant. A k-nearest-neighbours imputer averages the value from the most similar complete rows. An iterative imputer, inspired by the MICE method, predicts each incomplete column from the others, round after round. Some models skip filling altogether. scikit-learn’s histogram gradient boosting learns, at each split, which branch rows with a missing value should follow.
Filling is not free. Mean imputation squeezes a column’s spread. In an air-quality example, the ozone standard deviation fell from 33 to 28.7 after the blanks were filled with the mean. Predicting values from other columns looks smarter but makes the guesses too tidy, so correlations come out too strong. Because a guess is rarely as good as the real value, Google recommends a Boolean column that marks which values were imputed. scikit-learn’s imputers can add that flag through an add_indicator option. Statisticians warn that a related shortcut, the indicator method, which fills blanks with zero and adds a flag, can badly bias regression estimates. Imputers learn from data, so they belong after the train and test split. Fitting them on test rows is a form of data leakage.
Blanks in a dataset cause three problems.
Follow one column with two blanks through the options.
- 1 · splitSet the test rows aside before any fill value is computed.
- 2 · diagnoseCount the gaps in each column and ask why they are missing.
- 3 · fitLearn each column's fill value, such as its mean or median, from the training rows only.
- 4 · fillReplace each blank with that value and add a column that flags which values were filled.
- 5 · reuseApply the same stored values and flags to validation, test and live data.
Fill values come from training rows only, and a flag column tells the model which values were guessed.
| Who | What they ask | What it works with |
|---|---|---|
| Hospital analytics | “Are lab results blank because the test was never ordered?” | Empty lab-result cells and the reason each is empty |
| Survey team | “Do people with weaker opinions skip this question more often?” | Unanswered survey items |
| Streaming product team | “What does a watch time of minus one second mean?” | A magic value standing in for no measurement |
| Fraud team | “Can the model use the fact that an income field was left empty?” | An income_was_missing flag column |
- Models that need complete numeric input can train on data that had gaps.
- Imputation keeps rows that deletion would throw away.
- A flag column keeps the fact that a value was missing, which can itself be informative.
- Multiple imputation with MICE, as in statsmodels, gives standard errors that account for the gaps.
- Imputation does not recover the real values, and imputed ones are rarely as good.
- Mean imputation shrinks a column's spread and distorts its links to other columns.
- Most simple fixes are safe only when values are missing completely at random, which is often unrealistic.
- A single filled-in dataset hides how uncertain the guesses are, and scikit-learn's IterativeImputer returns just one.
Sources used
This explainer is written in original language. The links below support its factual claims.
- docsImputation of missing values, scikit-learn · read 27 Sept 2026
- docsSimpleImputer, scikit-learn · read 27 Sept 2026
- docsEnsembles: Gradient boosting, random forests, bagging, voting, stacking, scikit-learn · read 27 Sept 2026
- docsCommon pitfalls and recommended practices, scikit-learn · read 27 Sept 2026
- docsDatasets: Data characteristics (Machine Learning Crash Course), Google for Developers · read 27 Sept 2026
- docsNumerical data: Qualities of good numerical features (Machine Learning Crash Course), Google for Developers · read 27 Sept 2026
- docsConcepts of MCAR, MAR and MNAR (Flexible Imputation of Missing Data, section 1.2), Flexible Imputation of Missing Data, second edition (official online book) · read 27 Sept 2026
- docsAd-hoc solutions (Flexible Imputation of Missing Data, section 1.3), Flexible Imputation of Missing Data, second edition (official online book) · read 27 Sept 2026
- docsWorking with missing data, pandas · read 27 Sept 2026
- docsMultiple Imputation with Chained Equations, statsmodels · read 27 Sept 2026