Ridge, Lasso and Elastic Net
"Why does lasso set coefficients to exactly zero, and ridge does not?" The question has two answers, one from geometry and one from algebra, and a complete reply gives both. The regression lesson introduces both penalties. This lesson covers what ridge shrinks and why, the two answers, the Bayesian reading of both penalties, and the rules for using them on financial data.
Three penalties
Each method minimises the residual sum of squares plus a penalty on the coefficients , with a weight :
| Method | Penalty | Effect |
|---|---|---|
| Ridge | Shrinks every coefficient, none to zero | |
| Lasso | Shrinks, and sets some to exactly zero | |
| Elastic net | A weighted mix of the two | Selects, and shares weight across correlated features |
Two rules apply to all three. Standardise the features first, because the penalty treats a coefficient on a feature measured in basis points the same as one measured in per cent. And do not penalise the intercept.
The rest of this lesson is for subscribers
Unlock every lesson in Machine Learning for Quantitative Research, and every other premium course.
Subscribe to continueTest your knowledge
Keep reading Machine Learning for Quantitative Research
27 lessons in this course, and every other premium course, on one subscription.
- Every lesson in every course, with the worked examples and interactive simulators
- Graded questions on every lesson, with explanations for the wrong answers as well as the right one
- The trainers, timed assessments and brainteaser library that go with them