Bias–variance tradeoff
Bias is average error across training sets, while variance indicates sensitivity to varying training sets.
A model that is too simple can fit both the samples and the true function poorly. It can have high bias. A highly flexible model can fit the training data perfectly yet fail to fit the true function because it is sensitive to varying training data. It can have high variance.
For squared-error regression, expected error can be separated into squared bias, variance, and irreducible noise. A degree-1 polynomial can underfit when it is too simple for the training samples.
Model selection can balance loss against complexity. Cross-validation can be used directly for model selection. Bagging fitted trees can reduce variance enough to lower overall mean squared error.
Bias and variance describe two different contributors to generalization error.
Retrain two model families on three samples from the same process.
- 1 · resampleImagine drawing several training sets from the same underlying problem.
- 2 · refitFit the same learning method independently on each set.
- 3 · separateCompare the average prediction with the best possible model.
- 4 · chooseUse cross-validation to select a model and its settings.
Bias is average error across training sets; variance is sensitivity to varying training sets.
| Who | What they ask | What it works with |
|---|---|---|
| Risk modeller | “Is this scorecard consistently missing nonlinear patterns?” | Average residual and variation across resamples |
| Data scientist | “Does a small data change produce a very different tree?” | Predictions from repeated fitted samples |
| Model reviewer | “Did bagging reduce instability enough to lower error?” | Bias, variance, and error estimates |
- The decomposition gives separate names to systematic error, training-sample sensitivity, and irreducible noise in squared-error regression.
- Estimators have both bias and variance, so model selection tries to keep both low.
- It also explains how averaging fitted trees can reduce variance enough to lower total error.
- Validating a model still requires a scoring function.
- The decomposition does not remove the irreducible noise in the data.
- The stated decomposition applies to expected mean squared error in regression, not automatically to every metric.
Sources used
This explainer is written in original language. The links below support its factual claims.
- docsSingle estimator versus bagging: bias-variance decomposition, scikit-learn · read 28 Sept 2026
- docs3.5. Validation curves: plotting scores to evaluate models, scikit-learn · read 27 Sept 2026
- docsUnderfitting vs. Overfitting, scikit-learn · read 27 Sept 2026
- docsOverfitting: Model complexity, Google for Developers · read 27 Sept 2026
- docs3.1. Cross-validation: evaluating estimator performance, scikit-learn · read 27 Sept 2026