Courses

Machine Learning for Quantitative Research

AdvancedPremium8 sections · 27 lessons · 81 questionsLog in to track progress

Several systematic funds test machine learning in their researcher interviews, and some run a separate round for it. The round does not test whether you can call a library. It tests whether you know what a model assumes, why it fails, and how you would know that it had failed.

Financial data makes those questions hard. A daily return is mostly noise, the relationships drift, and the number of independent observations is much smaller than the row count. Methods that work well on images or text often fit noise here. A model that explains one per cent of the variance of returns out of sample can be worth a lot of money.

This course teaches machine learning with that problem in mind. It starts with what a model learns and why it overfits, derives the bias-variance decomposition, and shows how gradient descent fits a model. It then covers linear models: ridge, lasso and elastic net as penalties and as priors, logistic regression with the cross-entropy loss, and support vector machines with kernels. A section on evaluation follows: cross-validation and the choice of k, classification metrics and calibration, forecast metrics for returns such as the information coefficient, and hyperparameter tuning that does not fool you.

The second half covers tree ensembles, what feature importance does and does not mean, and how to combine several forecasts. It then turns to the data: how to build features, label returns, set cross-sectional targets, handle missing values and outliers, and retrain when the market drifts. PCA and clustering follow, then neural networks and sequence models. The last section puts it together: how to choose a model family on noisy data, and the machine learning round itself.

Each lesson starts from an interview question, works the numbers by hand, and ends with coding problems that you write from scratch in Python. The research validation course covers leakage, cross-validation on time-ordered data and multiple testing in full. This course covers those topics only as far as a model needs them, and links to that course for the rest.

Keep reading Machine Learning for Quantitative Research

27 lessons in this course, and every other premium course, on one subscription.

  • Every lesson in every course, with the worked examples and interactive simulators
  • Graded questions on every lesson, with explanations for the wrong answers as well as the right one
  • The trainers, timed assessments and brainteaser library that go with them