Ensemble learning
Ensemble learning combines the predictions of several models, usually by a vote or an average, so that their separate mistakes partly cancel out.
Ensemble learning means training several models and combining what they predict. For a category, such as spam or not spam, the members usually vote; for a number, such as a price, their answers are averaged. The members can be copies of one method trained on different samples of the data, or different kinds of model altogether. The best-known example is a random forest, which is built from many decision trees.
Why would a crowd of imperfect models beat each member? Each model must be better than random guessing, and the models must be wrong on different examples. In one simulated case, 21 models that are each wrong 30% of the time, and err independently, give a majority vote that is wrong only 2.6% of the time. Real models are rarely that independent, so the practical aim is errors that are at least somewhat uncorrelated. Under that equal-error, independent-error setup, voting gets worse when each classifier’s error exceeds 50 per cent. For regression with positive weights that sum to one and squared error, the ensemble’s error cannot exceed the members’ weighted average error. More disagreement, with members just as accurate, lowers it further. In one project to spot volcanoes on Venus, an ensemble of 32 neural networks matched the performance of human experts.
Bagging trains each member on a bootstrap sample, rows drawn at random with replacement, and a random forest gets its trees this way. Boosting builds members one after another, giving extra weight to the examples still being misclassified. Gradient boosting extends the idea to any differentiable loss. Stacking feeds the members’ predictions into a final model that learns how to combine them. In scikit-learn, that final model is logistic regression by default. It is trained on predictions made through cross-validation, to avoid overfitting. Plain voting simply combines different model types by majority or by averaged probabilities. With many wrong labels, boosting piles weight onto the mislabelled examples; in the paper’s experiments with noise, bagging still worked very well.
A single model has blind spots, and extra copies of it share every one.
Follow five test examples through three models and one vote.
- 1 · trainTrain several models that each beat random guessing and that disagree on some inputs.
- 2 · predictEach model makes its own prediction for a new example.
- 3 · combineThe ensemble takes a majority vote for a category, or an average for a number.
- 4 · weighOptionally, better models get more say, through weights or a final model trained to combine them.
A vote helps only when members are right more often than not and wrong in different places.
| Who | What they ask | What it works with |
|---|---|---|
| Email security team | “Is this message spam?” | A majority vote across many decision trees in a random forest |
| Card fraud team | “Should this payment be held for review?” | A tree model, a logistic regression and a neural network, voting |
| Retail planning | “How many umbrellas will this store sell next week?” | The average of several forecasting models |
| Imaging researchers | “Does this scan show the feature we are looking for?” | Several networks trained from different random starting weights |
- Accurate models whose errors differ can combine into an ensemble that beats every one of them.
- Bagging steadies unstable methods such as decision trees, whose output swings with small changes in the data.
- Boosting turns rough rules of thumb, each only a little better than chance, into a very accurate rule.
- The idea works on top of any model type, not only trees.
- Models that all make the same mistakes gain nothing from being combined.
- Bagging a stable method, such as nearest neighbours, barely changes it and can slightly hurt it.
- Stacking adds training cost, which scikit-learn describes as computationally expensive.
- A vote of many bagged trees loses the simple, readable structure of a single tree.
Sources used
This explainer is written in original language. The links below support its factual claims.
- paperEnsemble Methods in Machine Learning, Dietterich, Oregon State University · read 27 Sept 2026
- paperNeural Network Ensembles, Cross Validation, and Active Learning, Krogh and Vedelsby, NIPS 1994 (NeurIPS Proceedings) · read 27 Sept 2026
- paperBagging Predictors, Breiman, UC Berkeley Statistics (Technical Report 421) · read 27 Sept 2026
- paperA Short Introduction to Boosting, Freund and Schapire, AT&T Labs Research · read 27 Sept 2026
- docsEnsembles: Gradient boosting, random forests, bagging, voting, stacking, scikit-learn · read 27 Sept 2026
- docsStackingClassifier, scikit-learn · read 27 Sept 2026
- docsMachine Learning Glossary, Google for Developers · read 27 Sept 2026