Choosing a Model Family on Noisy Data
"Which models would you try on this problem, and why?" is the question the machine learning round most often opens with. Listing algorithms is a weak answer. A strong answer says what each family assumes about the signal, what it needs from the data, and how you would decide between them without fooling yourself.
What the three families assume
The families differ in what shape of signal they can represent, and in how much data they need before that flexibility helps rather than hurts.
Regularised linear models, such as ridge and lasso, assume each feature adds a fixed amount to the prediction. They are fast, stable and easy to inspect, and regularisation controls their variance directly.
The rest of this lesson is for subscribers
Unlock every lesson in Advanced Topics in Probability and Statistics, and every other premium course.
Subscribe to continueTest your knowledge
Keep reading Advanced Topics in Probability and Statistics
47 lessons in this course, and every other premium course, on one subscription.
- Every lesson in every course, with the worked examples and interactive simulators
- Graded questions on every lesson, with explanations for the wrong answers as well as the right one
- The trainers, timed assessments and brainteaser library that go with them