Boosting and Gradient Boosted Trees

"Bagging against boosting: what is the difference, and when does each fail?" is one of the most reliable questions in the machine learning round. Gradient boosted trees are also the model a researcher is most likely to have used on tabular data, so the follow-ups go into how they are tuned. The short answer to the first question: bagging averages many deep trees, grown independently, to reduce variance, and boosting adds many shallow trees in sequence to reduce bias.

The idea: fit what is still wrong

Boosting starts from a simple prediction, usually the mean of the target. Each round fits a small tree to the errors of the model so far, and adds a fraction of that tree's prediction to the model. Every new tree concentrates on the observations the ensemble still gets wrong.

The rest of this lesson is for subscribers

Unlock every lesson in Advanced Topics in Probability and Statistics, and every other premium course.

Subscribe to continue

Test your knowledge

Questions are only available to subscribers.

Keep reading Advanced Topics in Probability and Statistics

47 lessons in this course, and every other premium course, on one subscription.

  • Every lesson in every course, with the worked examples and interactive simulators
  • Graded questions on every lesson, with explanations for the wrong answers as well as the right one
  • The trainers, timed assessments and brainteaser library that go with them