Boosting and Gradient Boosted Trees

"Bagging against boosting: what is the difference, and when does each fail?" A complete answer covers what each method reduces, how each one can overfit, and how a boosted model is tuned. In short, bagging averages many deep trees, grown independently, to reduce variance. Boosting adds many shallow trees in sequence to reduce bias.

The idea: fit what is still wrong

Boosting starts from a simple prediction, usually the mean of the target. Each round fits a small tree to the errors of the model so far, and adds a fraction of that tree's prediction to the model. Every new tree concentrates on the observations the ensemble still gets wrong.

For squared-error loss, the errors to fit are the residuals ri=yi−Fm−1(xi)r_i = y_i - F_{m-1}(x_i), and round mm updates the model to

The rest of this lesson is for subscribers

Unlock every lesson in Machine Learning for Quantitative Research, and every other premium course.

Subscribe to continue

Test your knowledge

Questions are only available to subscribers.

Keep reading Machine Learning for Quantitative Research

27 lessons in this course, and every other premium course, on one subscription.

  • Every lesson in every course, with the worked examples and interactive simulators
  • Graded questions on every lesson, with explanations for the wrong answers as well as the right one
  • The trainers, timed assessments and brainteaser library that go with them