Statistics for machine learning
Statistics describes data and measures how far to trust a model's score, since every test set is only a sample of the cases the model will meet.
Statistics is the toolkit for describing data and for judging how far to trust what a sample shows. Before building features, Google’s Machine Learning Crash Course recommends computing basic statistics such as the mean, median and standard deviation. The mean is the sum of the values divided by how many there are. The median is the middle value, with half the data below it and half above, so extreme values pull the mean but not the median. Variance is roughly the average squared distance from the mean, and the standard deviation is its square root, back in the data’s own units. A link between two columns does not show that one causes the other: observational data shows association, and cause needs an experiment.
The second job is inference. A sample is a subset of a population, the whole group you care about, and a test set is a sample of the cases a model will meet. Its score is an estimate, not the true value. For a proportion such as accuracy, the standard error is the square root of p(1 − p)/n, where p is the score and n the number of test examples. A 95% confidence interval runs 1.96 standard errors either side. That does not mean a 95% chance that this interval holds the true value. It means that over many samples, about 95% of such intervals would. The rule needs independent examples and at least 10 successes and 10 failures. The Wilson method, a better interval, never dips below zero.
Comparing models is where the traps sit. For two models scored on the same examples, Dietterich’s 1998 study favoured McNemar’s test, which looks at the examples where they disagree. Cross-validation scores are not independent either, since models share the same folds, so scikit-learn’s example uses a corrected t-test. Running many tests invites a lucky winner, and the Bonferroni correction makes the threshold stricter to cut false positives. One benchmark study advises randomising seeds, data splits and other sources of variation. Finally, reusing a test set for many decisions wears it out, and a fair test uses new examples, not duplicates of training ones.
One score from one test set can mislead in three ways.
Follow one test score to an honest range.
- 1 · describeSummarise the data with a typical value, such as the mean or median, and a spread, such as the standard deviation.
- 2 · estimateTreat the test score as an estimate and compute its standard error from the score and the number of test examples.
- 3 · intervalBuild a 95% confidence interval that runs 1.96 standard errors either side of the score.
- 4 · compareCompare two models on the same examples with a paired test, such as McNemar's test.
- 5 · decideCall one model better only if the gap is both statistically significant and large enough to matter.
Fewer test examples mean a wider range.
| Who | What they ask | What it works with |
|---|---|---|
| Student comparing two classifiers | “Is my new model really better, or did I get a lucky test split?” | Confidence intervals and a paired test on one shared test set |
| Research team | “Does our half-point gain survive new random seeds and data splits?” | Scores repeated across seeds and splits |
| Data scientist cleaning a dataset | “Is something odd hiding in this column?” | Mean, median, standard deviation and percentiles |
| Team running a live model | “Has the incoming data changed since we trained?” | Serving data compared against training data |
- It summarises a column of data with a few numbers, such as its mean, median and standard deviation.
- A confidence interval shows how much uncertainty there is in a measured score.
- It explains why a larger test set gives a narrower, more precise range.
- Tests such as McNemar's give a way to check whether two classifiers really differ.
- An association in the data cannot prove cause and effect.
- A range measured on old test data says nothing about serving data that later changes.
- The simple normal-approximation interval can be inaccurate for very small samples or very few errors.
- Standard cross-validation assumes independent examples, which time-series data breaks.
Sources used
This explainer is written in original language. The links below support its factual claims.
- docsOpenIntro Statistics, Fourth Edition (screen-reader PDF), OpenIntro · read 27 Sept 2026
- official7.2.4.1. Confidence intervals (NIST/SEMATECH e-Handbook of Statistical Methods), National Institute of Standards and Technology · read 27 Sept 2026
- official1.3.5.2. Confidence Limits for the Mean (NIST/SEMATECH e-Handbook of Statistical Methods), National Institute of Standards and Technology · read 27 Sept 2026
- official1.3.5.1. Measures of Location (NIST/SEMATECH e-Handbook of Statistical Methods), National Institute of Standards and Technology · read 27 Sept 2026
- official1.3.5.6. Measures of Scale (NIST/SEMATECH e-Handbook of Statistical Methods), National Institute of Standards and Technology · read 27 Sept 2026
- official1.3.3.26. Scatter Plot (NIST/SEMATECH e-Handbook of Statistical Methods), National Institute of Standards and Technology · read 27 Sept 2026
- docsNumerical data: First steps (Machine Learning Crash Course), Google for Developers · read 27 Sept 2026
- docsDatasets: Dividing the original dataset (Machine Learning Crash Course), Google for Developers · read 27 Sept 2026
- docsProduction ML systems: Monitoring pipelines (Machine Learning Crash Course), Google for Developers · read 27 Sept 2026
- docs3.1. Cross-validation: evaluating estimator performance, scikit-learn developers · read 27 Sept 2026
- docsStatistical comparison of models using grid search, scikit-learn developers · read 27 Sept 2026
- paperModel Evaluation, Model Selection, and Algorithm Selection in Machine Learning, Raschka, arXiv 2018 · read 27 Sept 2026
- paperAccounting for Variance in Machine Learning Benchmarks, Bouthillier et al., arXiv 2021 · read 27 Sept 2026