Ridge, Lasso and Elastic Net

"Why does lasso set coefficients to exactly zero, and ridge does not?" The question has two answers, one from geometry and one from algebra, and a complete reply gives both. The regression lesson introduces both penalties. This lesson covers what ridge shrinks and why, the two answers, the Bayesian reading of both penalties, and the rules for using them on financial data.

Three penalties

Each method minimises the residual sum of squares plus a penalty on the coefficients β\beta, with a weight λ≥0\lambda \ge 0:

Method Penalty Effect
Ridge λ∑jβj2\lambda \sum_j \beta_j^2 Shrinks every coefficient, none to zero
Lasso λ∑j∣βj∣\lambda \sum_j \lvert \beta_j \rvert Shrinks, and sets some to exactly zero
Elastic net A weighted mix of the two Selects, and shares weight across correlated features

Two rules apply to all three. Standardise the features first, because the penalty treats a coefficient on a feature measured in basis points the same as one measured in per cent. And do not penalise the intercept.

The rest of this lesson is for subscribers

Unlock every lesson in Machine Learning for Quantitative Research, and every other premium course.

Subscribe to continue

Test your knowledge

Questions are only available to subscribers.

Keep reading Machine Learning for Quantitative Research

27 lessons in this course, and every other premium course, on one subscription.

  • Every lesson in every course, with the worked examples and interactive simulators
  • Graded questions on every lesson, with explanations for the wrong answers as well as the right one
  • The trainers, timed assessments and brainteaser library that go with them