Concepts

Feature importance

4 min readintermediateUpdated 28 Sept 2026
1 · In one line

Feature importance gives each input column a score for how much a trained model relies on it, so you can see what drives its predictions.

1 · What it is

Feature importance is a set of scores, one per input, showing how much a trained model depends on each. There are several ways to compute it, and each tells you something slightly different.

Tree ensembles such as random forests come with a built-in score called mean decrease in impurity, or MDI. When a tree splits its rows on a feature, the two groups it makes are less mixed than the group they came from. That mixing is called impurity, and the tree’s own split rule, such as Gini, is the yardstick for it. MDI credits each feature with all the mixing it cleared up, summed over every tree in the forest. It has two known flaws. First, it gives an unfair edge to columns that can take lots of values, like a price, over a simple yes-or-no column. A 2007 study in BMC Bioinformatics found the same bias toward variables with more categories. Second, MDI is worked out on the training data, so a feature the model used only to memorise noise can still score high. In a scikit-learn example on Titanic passenger data, a column of pure random numbers ranked among the forest’s most important features.

Permutation importance comes from Breiman’s 2001 random forest paper. He scrambled one variable in the rows a tree had not trained on and measured how much the error rate rose. scikit-learn applies the same idea to any fitted model on a held-out set. On the Titanic example, it put sex and passenger class on top and gave both random columns scores close to zero. Check the model is good first, because a feature that looks unimportant to a weak model could matter a lot to a strong one.

SHAP values, from a 2017 paper by Lundberg and Lee, work at the level of one prediction. They borrow Shapley values from cooperative game theory, treating features as players sharing credit for the model’s output. A row’s SHAP values add up to the gap between the model’s average output and its output for that row. Averaging their absolute size over many rows turns them back into an overall ranking.

None of these scores measures cause and effect. In a SHAP tutorial, a customer-renewal model credited bug reports with raising renewals, although reporting a bug had no causal effect. Correlated features blur the picture too, since shuffling one leaves its partner carrying the same signal. And many different models can fit the same data well while relying on different features.

2 · Why it exists

An accurate model gives answers, but on its own it does not say which inputs those answers depend on.

Opaque modelsThe most accurate models are often ensembles or deep networks, which even experts struggle to interpret.
Hidden data problemsSome problems in the data slip past standard accuracy checks, and seeing what the model relies on can expose them.
Extra inputsA model may carry features that add little, and removing the weak ones can make it more efficient.
3 · How it works

Follow permutation importance through one trained model and three features.

Illustrative numbers. Shuffling a column breaks its link to the target, and the drop in score is that feature's importance for this model.
  1. 1 · scoreMeasure the trained model's score, such as accuracy, on held-out data it did not train on.
  2. 2 · shuffleRandomly shuffle the values in one feature's column and leave every other column alone.
  3. 3 · rescoreRun the shuffled data through the same model and score it again.
  4. 4 · subtractThe feature's importance is the drop from the baseline score, averaged over several shuffles.
  5. 5 · rankRepeat for each feature; those whose shuffling hurts the score most rank highest.

The score says how much a feature matters to this model, not how useful the feature is in itself.

4 · Where it's used
WhoWhat they askWhat it works with
Credit risk team“Which inputs does our loan model lean on most?”Income, existing debt, account age and payment history
Hospital data scientists“Is the readmission model relying on a column it should never see?”Diagnosis codes, length of stay and discharge details
Subscription analytics“Which usage signals drive our renewal predictions?”Logins, bug reports and discounts
ML platform team“Which columns can we drop without losing accuracy?”The feature list of a production model
5 · What it solves, and what it doesn't
solves
  • Permutation importance works with any fitted model, not only trees.
  • Computed on held-out data, it shows which features help the model generalize.
  • Repeating the shuffle gives a spread of scores, not just one number.
  • SHAP values explain a single prediction, giving each feature its own share.
doesn't solve
  • It shows what a model relies on, not what causes the outcome in the real world.
  • When features are correlated, shuffling one leaves the other in place, so both can look unimportant.
  • Impurity-based scores favour features with many distinct values and can rank pure noise highly.
  • A feature that one good model ignores may be important to another good model.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. docsPermutation feature importance, scikit-learn · read 27 Sept 2026
  2. docsPermutation Importance vs Random Forest Feature Importance (MDI), scikit-learn · read 27 Sept 2026
  3. paperRandom Forests, Breiman, UC Berkeley (published in Machine Learning, 2001) · read 27 Sept 2026
  4. officialRandom forests - classification description, Breiman and Cutler, UC Berkeley Statistics · read 27 Sept 2026
  5. docsVariable importances (Decision Forests course), Google for Developers · read 27 Sept 2026
  6. paperA Unified Approach to Interpreting Model Predictions, Lundberg and Lee, NIPS 2017 (arXiv) · read 27 Sept 2026
  7. docsAn introduction to explainable AI with Shapley values, SHAP documentation · read 27 Sept 2026
  8. docsBe careful when interpreting predictive models in search of causal insights, SHAP documentation · read 27 Sept 2026
  9. docsIntroduction to Vertex Explainable AI, Google Cloud · read 27 Sept 2026
  10. paperBias in random forest variable importance measures: Illustrations, sources and a solution, Strobl et al., BMC Bioinformatics 2007 (PubMed Central) · read 27 Sept 2026
  11. paperAll Models are Wrong, but Many are Useful: Learning a Variable's Importance by Studying an Entire Class of Prediction Models Simultaneously, Fisher, Rudin and Dominici, Journal of Machine Learning Research · read 27 Sept 2026