Tuning Hyperparameters Without Fooling Yourself
"You tried 200 configurations and the best has a validation Sharpe ratio of 2. What do you report?" Not 2. That number is the maximum of 200 noisy estimates. A full answer works out what noise alone would produce. It then reports a score from data that the search never touched, together with the number of trials. This lesson covers the procedure behind that answer.
Three roles for data
- Training data fit the parameters: the weights, the splits, the coefficients.
- Validation data choose everything else: the hyperparameters, the model family, the features and the stopping point.
- Test data measure the final choice, once. See embargoes and the final holdout.
Every choice made on a set of data makes that set part of the fit for that choice. A validation score used to pick a winner is therefore no longer an unbiased estimate of the winner's error. On time-ordered data the three sets also come in time order, with purging between them where labels overlap. See walk-forward and purged cross-validation.
The rest of this lesson is for subscribers
Unlock every lesson in Machine Learning for Quantitative Research, and every other premium course.
Subscribe to continueTest your knowledge
Keep reading Machine Learning for Quantitative Research
27 lessons in this course, and every other premium course, on one subscription.
- Every lesson in every course, with the worked examples and interactive simulators
- Graded questions on every lesson, with explanations for the wrong answers as well as the right one
- The trainers, timed assessments and brainteaser library that go with them