Feature Importance and Its Traps
"Your random forest says feature X is the most important. What does that tell you?" A full answer first asks which measure was used, on which data, and what else is correlated with X. The ranking is a statement about one fitted model, and it can be wrong even about that model.
Impurity importance
The default in most tree libraries is mean decrease in impurity (MDI). Every split reduces the squared error or the Gini impurity by some amount. MDI credits that reduction to the feature the split used, sums it over all splits in all trees, and scales the totals to add up to one.
It has two problems.
- It is computed in sample. The reductions are measured on the data the trees were grown on. A deep tree that splits on noise still reduces training impurity, so a useless feature collects credit.
- It favours features with many split points. A continuous feature or a category with many levels offers hundreds of thresholds, and by chance one of them fits the training noise. A binary feature offers one threshold. A column of random numbers in a forest of deep trees usually earns a visible share of MDI.
The rest of this lesson is for subscribers
Unlock every lesson in Machine Learning for Quantitative Research, and every other premium course.
Subscribe to continueTest your knowledge
Keep reading Machine Learning for Quantitative Research
27 lessons in this course, and every other premium course, on one subscription.
- Every lesson in every course, with the worked examples and interactive simulators
- Graded questions on every lesson, with explanations for the wrong answers as well as the right one
- The trainers, timed assessments and brainteaser library that go with them