Shrinkage Estimators and Empirical Bayes

You have many noisy estimates: expected returns for 500 stocks, or the win rate of 50 strategies. The obvious approach estimates each independently.

Shrinkage does better: pull every estimate toward a common value.

Shrinkage
θ^ishrunk=λθˉ+(1λ)θ^i\hat{\theta}_i^{\text{shrunk}} = \lambda\bar{\theta} + (1-\lambda)\hat{\theta}_i

Pull each estimate toward the common mean. Biased on purpose, and it beats the unbiased estimator on total error.

A weighted blend of the individual estimate and the overall average, with λ\lambda controlling the strength.

The result that makes this surprising

James-Stein. When estimating three or more means simultaneously, the sample mean is inadmissible: a shrinkage estimator has lower total mean squared error, always, regardless of the true values.

This is genuinely counter-intuitive. It says that when estimating unrelated quantities, you should use information from all of them for each one.

The reason is that extreme sample values are extreme partly because of noise. The stock with the highest measured return probably had good luck as well as merit, so the best estimate of its true return is below what was measured. Shrinkage corrects that systematically.

The rest of this lesson is for subscribers

Unlock every lesson in Advanced Topics in Probability and Statistics, and every other premium course.

Subscribe to continue

Test your knowledge

Questions are only available to subscribers.

Keep reading Advanced Topics in Probability and Statistics

35 lessons in this course, and every other premium course, on one subscription.

  • Every lesson in every course, with the worked examples and interactive simulators
  • Graded questions on every lesson, with explanations for the wrong answers as well as the right one
  • The trainers, timed assessments and brainteaser library that go with them