Choosing a Model Family on Noisy Data

"Which models would you try on this problem, and why?" is the question the machine learning round most often opens with. Listing algorithms is a weak answer. A strong answer says what each family assumes about the signal, what it needs from the data, and how you would decide between them without fooling yourself.

What the three families assume

Regularised linear
CapturesAdditive, linear effects
NeedsLittle data, scaled features
Fails whenThe signal is non-linear
Explains itselfCoefficients
Tree ensembles
CapturesThresholds and interactions
NeedsModerate data, no scaling
Fails whenRelationships extrapolate or drift
Explains itselfImportance, partial dependence
Neural networks
CapturesAlmost any smooth shape
NeedsLarge data, careful tuning
Fails whenSamples are few and noisy
Explains itselfPoorly

The families differ in what shape of signal they can represent, and in how much data they need before that flexibility helps rather than hurts.

Regularised linear models, such as ridge and lasso, assume each feature adds a fixed amount to the prediction. They are fast, stable and easy to inspect, and regularisation controls their variance directly.

The rest of this lesson is for subscribers

Unlock every lesson in Advanced Topics in Probability and Statistics, and every other premium course.

Subscribe to continue

Test your knowledge

Questions are only available to subscribers.

Keep reading Advanced Topics in Probability and Statistics

47 lessons in this course, and every other premium course, on one subscription.

  • Every lesson in every course, with the worked examples and interactive simulators
  • Graded questions on every lesson, with explanations for the wrong answers as well as the right one
  • The trainers, timed assessments and brainteaser library that go with them