Combining Forecasts and Stacking

"You have two alpha signals with information coefficients of 0.04 and 0.03. How would you combine them?" A complete answer asks how correlated the signals are and works out how much the best mix could add. It then explains why equal weights are often the safer choice. The same reasoning covers stacking, where a second model learns how to combine the first ones.

Why averaging helps, and how much

Take MM forecasts whose errors each have variance σ2\sigma^2 and pairwise correlation ρ\rho. The average has error variance ρσ2+(1−ρ)σ2/M\rho\sigma^2 + (1 - \rho)\sigma^2 / M. This is the formula from decision trees and bagging, and it applies to any forecasts, not only trees.

The gain depends on the correlation. Two forecasts with error correlation 0.5 average to an error variance of 0.75σ20.75\sigma^2. With correlation 0.9, the average reaches only 0.95σ20.95\sigma^2. Two models trained on the same features with similar methods tend to make similar errors, so their average adds little.

The rest of this lesson is for subscribers

Unlock every lesson in Machine Learning for Quantitative Research, and every other premium course.

Subscribe to continue

Test your knowledge

Questions are only available to subscribers.

Keep reading Machine Learning for Quantitative Research

27 lessons in this course, and every other premium course, on one subscription.

  • Every lesson in every course, with the worked examples and interactive simulators
  • Graded questions on every lesson, with explanations for the wrong answers as well as the right one
  • The trainers, timed assessments and brainteaser library that go with them