Feature importance
Feature importance gives each input column a score for how much a trained model relies on it, so you can see what drives its predictions.
Feature importance is a set of scores, one per input, showing how much a trained model depends on each. There are several ways to compute it, and each tells you something slightly different.
Tree ensembles such as random forests come with a built-in score called mean decrease in impurity, or MDI. When a tree splits its rows on a feature, the two groups it makes are less mixed than the group they came from. That mixing is called impurity, and the tree’s own split rule, such as Gini, is the yardstick for it. MDI credits each feature with all the mixing it cleared up, summed over every tree in the forest. It has two known flaws. First, it gives an unfair edge to columns that can take lots of values, like a price, over a simple yes-or-no column. A 2007 study in BMC Bioinformatics found the same bias toward variables with more categories. Second, MDI is worked out on the training data, so a feature the model used only to memorise noise can still score high. In a scikit-learn example on Titanic passenger data, a column of pure random numbers ranked among the forest’s most important features.
Permutation importance comes from Breiman’s 2001 random forest paper. He scrambled one variable in the rows a tree had not trained on and measured how much the error rate rose. scikit-learn applies the same idea to any fitted model on a held-out set. On the Titanic example, it put sex and passenger class on top and gave both random columns scores close to zero. Check the model is good first, because a feature that looks unimportant to a weak model could matter a lot to a strong one.
SHAP values, from a 2017 paper by Lundberg and Lee, work at the level of one prediction. They borrow Shapley values from cooperative game theory, treating features as players sharing credit for the model’s output. A row’s SHAP values add up to the gap between the model’s average output and its output for that row. Averaging their absolute size over many rows turns them back into an overall ranking.
None of these scores measures cause and effect. In a SHAP tutorial, a customer-renewal model credited bug reports with raising renewals, although reporting a bug had no causal effect. Correlated features blur the picture too, since shuffling one leaves its partner carrying the same signal. And many different models can fit the same data well while relying on different features.
An accurate model gives answers, but on its own it does not say which inputs those answers depend on.
Follow permutation importance through one trained model and three features.
- 1 · scoreMeasure the trained model's score, such as accuracy, on held-out data it did not train on.
- 2 · shuffleRandomly shuffle the values in one feature's column and leave every other column alone.
- 3 · rescoreRun the shuffled data through the same model and score it again.
- 4 · subtractThe feature's importance is the drop from the baseline score, averaged over several shuffles.
- 5 · rankRepeat for each feature; those whose shuffling hurts the score most rank highest.
The score says how much a feature matters to this model, not how useful the feature is in itself.
| Who | What they ask | What it works with |
|---|---|---|
| Credit risk team | “Which inputs does our loan model lean on most?” | Income, existing debt, account age and payment history |
| Hospital data scientists | “Is the readmission model relying on a column it should never see?” | Diagnosis codes, length of stay and discharge details |
| Subscription analytics | “Which usage signals drive our renewal predictions?” | Logins, bug reports and discounts |
| ML platform team | “Which columns can we drop without losing accuracy?” | The feature list of a production model |
- Permutation importance works with any fitted model, not only trees.
- Computed on held-out data, it shows which features help the model generalize.
- Repeating the shuffle gives a spread of scores, not just one number.
- SHAP values explain a single prediction, giving each feature its own share.
- It shows what a model relies on, not what causes the outcome in the real world.
- When features are correlated, shuffling one leaves the other in place, so both can look unimportant.
- Impurity-based scores favour features with many distinct values and can rank pure noise highly.
- A feature that one good model ignores may be important to another good model.
Sources used
This explainer is written in original language. The links below support its factual claims.
- docsPermutation feature importance, scikit-learn · read 27 Sept 2026
- docsPermutation Importance vs Random Forest Feature Importance (MDI), scikit-learn · read 27 Sept 2026
- paperRandom Forests, Breiman, UC Berkeley (published in Machine Learning, 2001) · read 27 Sept 2026
- officialRandom forests - classification description, Breiman and Cutler, UC Berkeley Statistics · read 27 Sept 2026
- docsVariable importances (Decision Forests course), Google for Developers · read 27 Sept 2026
- paperA Unified Approach to Interpreting Model Predictions, Lundberg and Lee, NIPS 2017 (arXiv) · read 27 Sept 2026
- docsAn introduction to explainable AI with Shapley values, SHAP documentation · read 27 Sept 2026
- docsBe careful when interpreting predictive models in search of causal insights, SHAP documentation · read 27 Sept 2026
- docsIntroduction to Vertex Explainable AI, Google Cloud · read 27 Sept 2026
- paperBias in random forest variable importance measures: Illustrations, sources and a solution, Strobl et al., BMC Bioinformatics 2007 (PubMed Central) · read 27 Sept 2026
- paperAll Models are Wrong, but Many are Useful: Learning a Variable's Importance by Studying an Entire Class of Prediction Models Simultaneously, Fisher, Rudin and Dominici, Journal of Machine Learning Research · read 27 Sept 2026