Gradient boosting
Gradient boosting builds one strong model from many small decision trees, added one at a time, each trained to fix the errors the trees before it still make.
Gradient boosting is a way to build one strong model out of many weak ones, usually small decision trees. The trees are added in sequence. Each new tree is trained to predict what the model so far still gets wrong, and its output is added to the running prediction. The method was set out in a 2001 paper in The Annals of Statistics. That paper treated learning as optimisation in the space of functions, not parameters, and built a general gradient descent boosting method for any fitting criterion.
The name comes from that link to gradient descent. For squared error, the thing each tree fits is simply true value minus prediction, called the residual. For other losses, each tree fits the negative gradient: the direction in which each example’s prediction should move to lower the loss. Learning all the trees at once is intractable, so they are added one at a time while earlier ones stay fixed. Each tree’s output is multiplied by a learning rate, also called shrinkage, before it is added. A common learning rate is 0.1, and values nearer 0 reduce overfitting more than values nearer 1.
A random forest, by contrast, averages trees trained with added randomness, so some errors cancel out. Boosted trees are more prone to overfitting than a forest, so they need extra controls. Common ones are a maximum tree depth and the share of features tested at each node. Stochastic gradient boosting trains each tree on a random fraction of the training data. scikit-learn’s histogram-based estimators turn on early stopping by default from 10,000 training samples, watching the loss on held-back validation data.
XGBoost’s paper describes a system that scales beyond billions of examples. LightGBM’s authors report training up to over 20 times faster than conventional gradient boosting with almost the same accuracy. CatBoost adds ordered boosting and a new way to handle categorical features, and can stop adding trees when its overfitting detector triggers. scikit-learn’s HistGradientBoosting estimators, inspired by LightGBM, can be orders of magnitude faster on tens of thousands of samples or more.
One small tree is too crude, one deep tree overfits, and independent trees never aim at each other's mistakes.
Follow three house prices through one round of boosting.
- 1 · startThe model begins with one constant guess, which for squared error is the average of the training targets.
- 2 · measureIt works out each example's error, which for squared error is the gap between the true value and the prediction.
- 3 · fitA small tree is trained to predict those errors, or more generally the negative gradient of the loss.
- 4 · addThe tree's output, scaled down by the learning rate, is added to the running prediction.
- 5 · repeatThe loop runs until a set number of trees is reached or the model starts to overfit on validation data.
The finished model is an additive model: the starting guess plus every scaled tree.
| Who | What they ask | What it works with |
|---|---|---|
| Credit risk team | “Which loan applications are most likely to default?” | Income, existing debt and repayment history |
| Retail planning | “How many units will each store sell next week?” | Past sales, promotions and season |
| Fraud team | “Which card payments look unusual for this customer?” | Amount, merchant and time of day |
| Insurance pricing | “What is this policy likely to cost in claims?” | Vehicle, driver history and past claims |
- It is a strong choice for regression and classification on tabular data, meaning rows and columns like a spreadsheet.
- Trained models are small and fast, often taking a few microseconds per prediction.
- Numerical and categorical features usually work with little preprocessing.
- It works with any differentiable loss, so one method covers many kinds of task.
- Unlike a random forest, it can overfit, so it needs limits and early stopping on validation data.
- Trees must be trained one after another, which can slow training considerably.
- On images and text it often does worse than other methods, because trees cannot learn and reuse internal representations.
- Its settings interact, since a smaller learning rate needs more trees to reach the same training error.
Sources used
This explainer is written in original language. The links below support its factual claims.
- paperGreedy function approximation: A gradient boosting machine, The Annals of Statistics (Institute of Mathematical Statistics), 2001 · read 27 Sept 2026
- docsGradient Boosted Decision Trees (Decision Forests course), Google for Developers · read 27 Sept 2026
- docsOverfitting, regularization, and early stopping (Decision Forests course), Google for Developers · read 27 Sept 2026
- docsRandom forests (Decision Forests course), Google for Developers · read 27 Sept 2026
- docsEnsembles: Gradient boosting, random forests, bagging, voting, stacking, scikit-learn · read 27 Sept 2026
- docsIntroduction to Boosted Trees, XGBoost developers · read 27 Sept 2026
- paperXGBoost: A Scalable Tree Boosting System, Chen and Guestrin, arXiv (KDD 2016) · read 27 Sept 2026
- paperLightGBM: A Highly Efficient Gradient Boosting Decision Tree, Ke et al., NeurIPS 2017 · read 27 Sept 2026
- docsHow training is performed, CatBoost · read 27 Sept 2026
- paperCatBoost: unbiased boosting with categorical features, Prokhorenkova et al., arXiv · read 27 Sept 2026