Decision Trees and Bagging

"Explain bagging, and why a random forest samples features" is a common question in the machine learning round, and so is its follow-up: "when would it mislead you on financial data?" Both answers come from one idea. A single deep tree has low bias and high variance, and averaging many trees removes variance, but only as far as the trees disagree.

How a tree splits

A regression tree divides the data with a sequence of yes-or-no questions of the form "is feature jj below threshold tt?". At each node it tries every feature and every threshold, and keeps the split that most reduces the squared error within the two resulting groups. A leaf predicts the mean of the training targets that reach it. Classification trees do the same with an impurity measure, such as Gini impurity or entropy, in place of squared error.

The rest of this lesson is for subscribers

Unlock every lesson in Advanced Topics in Probability and Statistics, and every other premium course.

Subscribe to continue

Test your knowledge

Questions are only available to subscribers.

Keep reading Advanced Topics in Probability and Statistics

47 lessons in this course, and every other premium course, on one subscription.

  • Every lesson in every course, with the worked examples and interactive simulators
  • Graded questions on every lesson, with explanations for the wrong answers as well as the right one
  • The trainers, timed assessments and brainteaser library that go with them