Concepts

Missing data

5 min readintermediateUpdated 28 Sept 2026
1 · In one line

Missing data are empty cells in a dataset. You can drop them, fill them with reasoned guesses or use a model that accepts gaps, and why they are missing decides which is safe.

1 · What it is

Missing data are cells in a dataset that hold no value. Real datasets often have them, stored as blanks, NaN (short for “not a number”) or another placeholder. In pandas, the marker for a missing value, called NA, depends on the column’s data type. A made-up code such as minus one for “no measurement” is worse, because the model treats it as a real number.

How the gaps arose matters as much as how many there are. Rubin sorted missing data into three categories in 1976. Missing completely at random, or MCAR, means every value had the same chance of going missing, so the gaps are unrelated to the data. A weighing scale whose battery ran out is the classic example. Missing at random, or MAR, means the chance differs only between groups you can see in the data. A scale that fails more often on soft floors fits, as long as the floor type was recorded. Missing not at random, or MNAR, means the chance depends on reasons you cannot see, such as people with weaker opinions answering a poll less often. Most modern methods assume at least MAR.

Deleting rows is the bluntest option, and under MCAR it still gives unbiased means and regression weights. In pandas, dropna removes rows or columns that contain gaps. Dropping a whole column makes sense when one feature is mostly empty and adds little. Google’s crash course suggests building one dataset each way and keeping whichever trains the better model.

Imputation means filling the gaps with reasoned guesses. Simple imputation fills each blank with a column’s mean, median or most frequent value, or with a fixed constant. A k-nearest-neighbours imputer averages the value from the most similar complete rows. An iterative imputer, inspired by the MICE method, predicts each incomplete column from the others, round after round. Some models skip filling altogether. scikit-learn’s histogram gradient boosting learns, at each split, which branch rows with a missing value should follow.

Filling is not free. Mean imputation squeezes a column’s spread. In an air-quality example, the ozone standard deviation fell from 33 to 28.7 after the blanks were filled with the mean. Predicting values from other columns looks smarter but makes the guesses too tidy, so correlations come out too strong. Because a guess is rarely as good as the real value, Google recommends a Boolean column that marks which values were imputed. scikit-learn’s imputers can add that flag through an add_indicator option. Statisticians warn that a related shortcut, the indicator method, which fills blanks with zero and adds a flag, can badly bias regression estimates. Imputers learn from data, so they belong after the train and test split. Fitting them on test rows is a form of data leakage.

2 · Why it exists

Blanks in a dataset cause three problems.

Models reject gapsMany scikit-learn models expect every cell to hold a meaningful number, so a blank or NaN stops them working.
Deleting wastes dataDropping every incomplete row is easy, but with many columns it can throw away more than half of the rows.
Guesses can misleadIf the gaps are not random, a quick fix can bias the averages and relationships the model learns.
3 · How it works

Follow one column with two blanks through the options.

Illustrative numbers. The fill value 40 is the mean of the four known training ages, and later rows reuse it.
  1. 1 · splitSet the test rows aside before any fill value is computed.
  2. 2 · diagnoseCount the gaps in each column and ask why they are missing.
  3. 3 · fitLearn each column's fill value, such as its mean or median, from the training rows only.
  4. 4 · fillReplace each blank with that value and add a column that flags which values were filled.
  5. 5 · reuseApply the same stored values and flags to validation, test and live data.

Fill values come from training rows only, and a flag column tells the model which values were guessed.

4 · Where it's used
WhoWhat they askWhat it works with
Hospital analytics“Are lab results blank because the test was never ordered?”Empty lab-result cells and the reason each is empty
Survey team“Do people with weaker opinions skip this question more often?”Unanswered survey items
Streaming product team“What does a watch time of minus one second mean?”A magic value standing in for no measurement
Fraud team“Can the model use the fact that an income field was left empty?”An income_was_missing flag column
5 · What it solves, and what it doesn't
solves
  • Models that need complete numeric input can train on data that had gaps.
  • Imputation keeps rows that deletion would throw away.
  • A flag column keeps the fact that a value was missing, which can itself be informative.
  • Multiple imputation with MICE, as in statsmodels, gives standard errors that account for the gaps.
doesn't solve
  • Imputation does not recover the real values, and imputed ones are rarely as good.
  • Mean imputation shrinks a column's spread and distorts its links to other columns.
  • Most simple fixes are safe only when values are missing completely at random, which is often unrealistic.
  • A single filled-in dataset hides how uncertain the guesses are, and scikit-learn's IterativeImputer returns just one.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. docsImputation of missing values, scikit-learn · read 27 Sept 2026
  2. docsSimpleImputer, scikit-learn · read 27 Sept 2026
  3. docsEnsembles: Gradient boosting, random forests, bagging, voting, stacking, scikit-learn · read 27 Sept 2026
  4. docsCommon pitfalls and recommended practices, scikit-learn · read 27 Sept 2026
  5. docsDatasets: Data characteristics (Machine Learning Crash Course), Google for Developers · read 27 Sept 2026
  6. docsNumerical data: Qualities of good numerical features (Machine Learning Crash Course), Google for Developers · read 27 Sept 2026
  7. docsConcepts of MCAR, MAR and MNAR (Flexible Imputation of Missing Data, section 1.2), Flexible Imputation of Missing Data, second edition (official online book) · read 27 Sept 2026
  8. docsAd-hoc solutions (Flexible Imputation of Missing Data, section 1.3), Flexible Imputation of Missing Data, second edition (official online book) · read 27 Sept 2026
  9. docsWorking with missing data, pandas · read 27 Sept 2026
  10. docsMultiple Imputation with Chained Equations, statsmodels · read 27 Sept 2026