Concepts

Gradient boosting

4 min readintermediateUpdated 28 Sept 2026
1 · In one line

Gradient boosting builds one strong model from many small decision trees, added one at a time, each trained to fix the errors the trees before it still make.

1 · What it is

Gradient boosting is a way to build one strong model out of many weak ones, usually small decision trees. The trees are added in sequence. Each new tree is trained to predict what the model so far still gets wrong, and its output is added to the running prediction. The method was set out in a 2001 paper in The Annals of Statistics. That paper treated learning as optimisation in the space of functions, not parameters, and built a general gradient descent boosting method for any fitting criterion.

The name comes from that link to gradient descent. For squared error, the thing each tree fits is simply true value minus prediction, called the residual. For other losses, each tree fits the negative gradient: the direction in which each example’s prediction should move to lower the loss. Learning all the trees at once is intractable, so they are added one at a time while earlier ones stay fixed. Each tree’s output is multiplied by a learning rate, also called shrinkage, before it is added. A common learning rate is 0.1, and values nearer 0 reduce overfitting more than values nearer 1.

A random forest, by contrast, averages trees trained with added randomness, so some errors cancel out. Boosted trees are more prone to overfitting than a forest, so they need extra controls. Common ones are a maximum tree depth and the share of features tested at each node. Stochastic gradient boosting trains each tree on a random fraction of the training data. scikit-learn’s histogram-based estimators turn on early stopping by default from 10,000 training samples, watching the loss on held-back validation data.

XGBoost’s paper describes a system that scales beyond billions of examples. LightGBM’s authors report training up to over 20 times faster than conventional gradient boosting with almost the same accuracy. CatBoost adds ordered boosting and a new way to handle categorical features, and can stop adding trees when its overfitting detector triggers. scikit-learn’s HistGradientBoosting estimators, inspired by LightGBM, can be orders of magnitude faster on tens of thousands of samples or more.

2 · Why it exists

One small tree is too crude, one deep tree overfits, and independent trees never aim at each other's mistakes.

Too weak aloneA shallow tree learns only a coarse outline of the pattern. Boosting combines many of these weak models into a strong one.
Deep trees overfitA deep tree fits its own training examples too closely. Boosting uses shallow trees, which underfit on their own, and adds more of them.
Trees that ignore each otherIn a random forest each tree is built independently of the rest. In boosting, each new tree is trained to correct the ones before it.
3 · How it works

Follow three house prices through one round of boosting.

Illustrative numbers. Each round fits a small tree to what is still wrong, then adds only half of that tree's output.
  1. 1 · startThe model begins with one constant guess, which for squared error is the average of the training targets.
  2. 2 · measureIt works out each example's error, which for squared error is the gap between the true value and the prediction.
  3. 3 · fitA small tree is trained to predict those errors, or more generally the negative gradient of the loss.
  4. 4 · addThe tree's output, scaled down by the learning rate, is added to the running prediction.
  5. 5 · repeatThe loop runs until a set number of trees is reached or the model starts to overfit on validation data.

The finished model is an additive model: the starting guess plus every scaled tree.

4 · Where it's used
WhoWhat they askWhat it works with
Credit risk team“Which loan applications are most likely to default?”Income, existing debt and repayment history
Retail planning“How many units will each store sell next week?”Past sales, promotions and season
Fraud team“Which card payments look unusual for this customer?”Amount, merchant and time of day
Insurance pricing“What is this policy likely to cost in claims?”Vehicle, driver history and past claims
5 · What it solves, and what it doesn't
solves
  • It is a strong choice for regression and classification on tabular data, meaning rows and columns like a spreadsheet.
  • Trained models are small and fast, often taking a few microseconds per prediction.
  • Numerical and categorical features usually work with little preprocessing.
  • It works with any differentiable loss, so one method covers many kinds of task.
doesn't solve
  • Unlike a random forest, it can overfit, so it needs limits and early stopping on validation data.
  • Trees must be trained one after another, which can slow training considerably.
  • On images and text it often does worse than other methods, because trees cannot learn and reuse internal representations.
  • Its settings interact, since a smaller learning rate needs more trees to reach the same training error.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. paperGreedy function approximation: A gradient boosting machine, The Annals of Statistics (Institute of Mathematical Statistics), 2001 · read 27 Sept 2026
  2. docsGradient Boosted Decision Trees (Decision Forests course), Google for Developers · read 27 Sept 2026
  3. docsOverfitting, regularization, and early stopping (Decision Forests course), Google for Developers · read 27 Sept 2026
  4. docsRandom forests (Decision Forests course), Google for Developers · read 27 Sept 2026
  5. docsEnsembles: Gradient boosting, random forests, bagging, voting, stacking, scikit-learn · read 27 Sept 2026
  6. docsIntroduction to Boosted Trees, XGBoost developers · read 27 Sept 2026
  7. paperXGBoost: A Scalable Tree Boosting System, Chen and Guestrin, arXiv (KDD 2016) · read 27 Sept 2026
  8. paperLightGBM: A Highly Efficient Gradient Boosting Decision Tree, Ke et al., NeurIPS 2017 · read 27 Sept 2026
  9. docsHow training is performed, CatBoost · read 27 Sept 2026
  10. paperCatBoost: unbiased boosting with categorical features, Prokhorenkova et al., arXiv · read 27 Sept 2026