Concepts

Bias–variance tradeoff

4 min readintermediateUpdated 28 Sept 2026
1 · In one line

Bias is average error across training sets, while variance indicates sensitivity to varying training sets.

1 · What it is

A model that is too simple can fit both the samples and the true function poorly. It can have high bias. A highly flexible model can fit the training data perfectly yet fail to fit the true function because it is sensitive to varying training data. It can have high variance.

For squared-error regression, expected error can be separated into squared bias, variance, and irreducible noise. A degree-1 polynomial can underfit when it is too simple for the training samples.

Model selection can balance loss against complexity. Cross-validation can be used directly for model selection. Bagging fitted trees can reduce variance enough to lower overall mean squared error.

2 · Why it exists

Bias and variance describe two different contributors to generalization error.

Systematic missBias measures how far the model's average prediction is from the best possible prediction.
Sample sensitivityVariance measures how much predictions change when the model is fitted on different samples from the same problem.
Unavoidable noiseSome error comes from variability in the data rather than the fitted model.
3 · How it works

Retrain two model families on three samples from the same process.

Retraining separates average miss from sample sensitivity Three samples feed rigid and flexible model families. Rigid predictions cluster away from the target; flexible predictions spread around it. Validation compares total error. THREE SAMPLES A · B · C same process Rigid family refit on A · B · C similar fitted models Flexible family refit on A · B · C different fitted models PREDICTIONS target: 50 models: 31 · 32 · 30 high bias · low variance PREDICTIONS target: 50 models: 29 · 51 · 70 low bias · high variance KEY STEP Validate compare total error Average offset measures bias; movement across refits measures variance.
Retraining shows how sensitive predictions are to changes in the training set.
  1. 1 · resampleImagine drawing several training sets from the same underlying problem.
  2. 2 · refitFit the same learning method independently on each set.
  3. 3 · separateCompare the average prediction with the best possible model.
  4. 4 · chooseUse cross-validation to select a model and its settings.

Bias is average error across training sets; variance is sensitivity to varying training sets.

4 · Where it's used
WhoWhat they askWhat it works with
Risk modeller“Is this scorecard consistently missing nonlinear patterns?”Average residual and variation across resamples
Data scientist“Does a small data change produce a very different tree?”Predictions from repeated fitted samples
Model reviewer“Did bagging reduce instability enough to lower error?”Bias, variance, and error estimates
5 · What it solves, and what it doesn't
solves
  • The decomposition gives separate names to systematic error, training-sample sensitivity, and irreducible noise in squared-error regression.
  • Estimators have both bias and variance, so model selection tries to keep both low.
  • It also explains how averaging fitted trees can reduce variance enough to lower total error.
doesn't solve
  • Validating a model still requires a scoring function.
  • The decomposition does not remove the irreducible noise in the data.
  • The stated decomposition applies to expected mean squared error in regression, not automatically to every metric.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. docsSingle estimator versus bagging: bias-variance decomposition, scikit-learn · read 28 Sept 2026
  2. docs3.5. Validation curves: plotting scores to evaluate models, scikit-learn · read 27 Sept 2026
  3. docsUnderfitting vs. Overfitting, scikit-learn · read 27 Sept 2026
  4. docsOverfitting: Model complexity, Google for Developers · read 27 Sept 2026
  5. docs3.1. Cross-validation: evaluating estimator performance, scikit-learn · read 27 Sept 2026