Statistics for MLConcepts

Statistics for machine learning

4 min readbeginnerUpdated 28 Sept 2026
1 · In one line

Statistics describes data and measures how far to trust a model's score, since every test set is only a sample of the cases the model will meet.

1 · What it is

Statistics is the toolkit for describing data and for judging how far to trust what a sample shows. Before building features, Google’s Machine Learning Crash Course recommends computing basic statistics such as the mean, median and standard deviation. The mean is the sum of the values divided by how many there are. The median is the middle value, with half the data below it and half above, so extreme values pull the mean but not the median. Variance is roughly the average squared distance from the mean, and the standard deviation is its square root, back in the data’s own units. A link between two columns does not show that one causes the other: observational data shows association, and cause needs an experiment.

The second job is inference. A sample is a subset of a population, the whole group you care about, and a test set is a sample of the cases a model will meet. Its score is an estimate, not the true value. For a proportion such as accuracy, the standard error is the square root of p(1 − p)/n, where p is the score and n the number of test examples. A 95% confidence interval runs 1.96 standard errors either side. That does not mean a 95% chance that this interval holds the true value. It means that over many samples, about 95% of such intervals would. The rule needs independent examples and at least 10 successes and 10 failures. The Wilson method, a better interval, never dips below zero.

Comparing models is where the traps sit. For two models scored on the same examples, Dietterich’s 1998 study favoured McNemar’s test, which looks at the examples where they disagree. Cross-validation scores are not independent either, since models share the same folds, so scikit-learn’s example uses a corrected t-test. Running many tests invites a lucky winner, and the Bonferroni correction makes the threshold stricter to cut false positives. One benchmark study advises randomising seeds, data splits and other sources of variation. Finally, reusing a test set for many decisions wears it out, and a fair test uses new examples, not duplicates of training ones.

2 · Why it exists

One score from one test set can mislead in three ways.

Samples wobbleA test set is a sample, so its score is an estimate. The smaller the test set, the more that estimate varies.
Small gaps look realThe uncertainty from which examples happened to be sampled is not small compared with typical improvements between models.
Many tries find luckA significance threshold that suits one comparison is not suitable when many pairs of models are compared.
3 · How it works

Follow one test score to an honest range.

Illustrative numbers. A 95% range is 1.96 standard errors either side of the score, and it shrinks as the test set grows.
  1. 1 · describeSummarise the data with a typical value, such as the mean or median, and a spread, such as the standard deviation.
  2. 2 · estimateTreat the test score as an estimate and compute its standard error from the score and the number of test examples.
  3. 3 · intervalBuild a 95% confidence interval that runs 1.96 standard errors either side of the score.
  4. 4 · compareCompare two models on the same examples with a paired test, such as McNemar's test.
  5. 5 · decideCall one model better only if the gap is both statistically significant and large enough to matter.

Fewer test examples mean a wider range.

4 · Where it's used
WhoWhat they askWhat it works with
Student comparing two classifiers“Is my new model really better, or did I get a lucky test split?”Confidence intervals and a paired test on one shared test set
Research team“Does our half-point gain survive new random seeds and data splits?”Scores repeated across seeds and splits
Data scientist cleaning a dataset“Is something odd hiding in this column?”Mean, median, standard deviation and percentiles
Team running a live model“Has the incoming data changed since we trained?”Serving data compared against training data
5 · What it solves, and what it doesn't
solves
  • It summarises a column of data with a few numbers, such as its mean, median and standard deviation.
  • A confidence interval shows how much uncertainty there is in a measured score.
  • It explains why a larger test set gives a narrower, more precise range.
  • Tests such as McNemar's give a way to check whether two classifiers really differ.
doesn't solve
  • An association in the data cannot prove cause and effect.
  • A range measured on old test data says nothing about serving data that later changes.
  • The simple normal-approximation interval can be inaccurate for very small samples or very few errors.
  • Standard cross-validation assumes independent examples, which time-series data breaks.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. docsOpenIntro Statistics, Fourth Edition (screen-reader PDF), OpenIntro · read 27 Sept 2026
  2. official7.2.4.1. Confidence intervals (NIST/SEMATECH e-Handbook of Statistical Methods), National Institute of Standards and Technology · read 27 Sept 2026
  3. official1.3.5.2. Confidence Limits for the Mean (NIST/SEMATECH e-Handbook of Statistical Methods), National Institute of Standards and Technology · read 27 Sept 2026
  4. official1.3.5.1. Measures of Location (NIST/SEMATECH e-Handbook of Statistical Methods), National Institute of Standards and Technology · read 27 Sept 2026
  5. official1.3.5.6. Measures of Scale (NIST/SEMATECH e-Handbook of Statistical Methods), National Institute of Standards and Technology · read 27 Sept 2026
  6. official1.3.3.26. Scatter Plot (NIST/SEMATECH e-Handbook of Statistical Methods), National Institute of Standards and Technology · read 27 Sept 2026
  7. docsNumerical data: First steps (Machine Learning Crash Course), Google for Developers · read 27 Sept 2026
  8. docsDatasets: Dividing the original dataset (Machine Learning Crash Course), Google for Developers · read 27 Sept 2026
  9. docsProduction ML systems: Monitoring pipelines (Machine Learning Crash Course), Google for Developers · read 27 Sept 2026
  10. docs3.1. Cross-validation: evaluating estimator performance, scikit-learn developers · read 27 Sept 2026
  11. docsStatistical comparison of models using grid search, scikit-learn developers · read 27 Sept 2026
  12. paperModel Evaluation, Model Selection, and Algorithm Selection in Machine Learning, Raschka, arXiv 2018 · read 27 Sept 2026
  13. paperAccounting for Variance in Machine Learning Benchmarks, Bouthillier et al., arXiv 2021 · read 27 Sept 2026