Free preview
Overfitting and the Train-Test Gap
A typical interview question is: "What is overfitting, and how would you know if your model is doing it?" A full answer has four parts. It gives the definition and the diagnosis. It explains why financial data makes overfitting so easy. And it gives an example from your own work.
The definition
A model overfits when it fits the noise in its training sample as well as the signal. It then performs worse on new data than its fit on the training data suggests. A more useful working version: a model overfits when a simpler model would do better out of sample.
Every sample contains both signal, which repeats in new data, and noise, which does not. A flexible enough model can fit both. The part of the fit that came from noise is lost the moment the model meets new data.
Diagnose it from two numbers
Compare the error on the data the model was trained on with the error on data it has not seen.
| Training error | Validation error | Diagnosis |
|---|---|---|
| Low | Much higher | Overfitting: high variance |
| High | Close to training | Underfitting: high bias |
| Low | Close to training | A model that generalises |
On returns, the same comparison appears as an in-sample Sharpe ratio of 3 that falls to 0.4 out of sample, or a training of 0.40 against 0.02 on validation. The gap between the two numbers is the size of the overfit.
A learning curve adds a second view: plot both errors as the training set grows. If the gap narrows as data is added, the model has high variance and more data will help. If both errors level off high and close together, the model is too simple and more data will not help.
Bias, variance and noise
For squared error, the expected error of a model's prediction on new data splits into three parts:
Bias is how far the model's average prediction is from the truth. Variance is how much the prediction changes from one training sample to another. is the noise that no model can remove. A more flexible model lowers bias and raises variance. The next lesson, the bias-variance decomposition, derives this split and works it through by hand.
In most prediction problems outside finance, the noise term is small and flexibility pays. When predicting returns, the noise term dominates: a daily return is mostly noise, and a model with a small out-of-sample can be valuable. With so little signal, the variance that flexibility adds usually costs more than the bias it removes. That is the main reason simple models are hard to beat on financial data.
Many features, few observations
Overfitting is guaranteed when features are plentiful relative to observations. Fit least squares with features of pure noise, plus an intercept, to observations of a target that is also noise. The expected in-sample is
With random features and observations, the expected in-sample is . Half the variance of a target that has no relationship with any feature appears to be explained. With , the regression fits the training data perfectly.
Out of sample, the same model's is below zero. Forecast metrics for returns defines that measure and the out-of-sample Sharpe ratio. The predictions vary, but not with the target, so they add error to simply predicting the mean.
Selection makes it worse. Suppose a research process tries 100 noise features and keeps the 50 with the largest absolute correlation with the target. The in-sample of the refit is then about 0.65, not 0.5. That is the average over 2,000 simulated runs with . On 1,000 new observations of the same noise, the was below zero in every run.
The noise benchmark
The most convincing test of a whole research pipeline is to run it on data that contains no signal. Shuffle the target, or replace it with random numbers, and run the same feature selection, tuning and validation. If the pipeline still reports good out-of-sample results, something in it is leaking information or selecting on noise. The noise benchmark catches problems that no single statistic does, because it tests the process rather than one model.
What reduces overfitting
- Fewer and better features, chosen before looking at the results.
- Regularisation, which shrinks coefficients towards zero; see ridge, lasso and elastic net and shrinkage estimators.
- More data, when the learning curve shows high variance.
- Validation that respects time, such as walk-forward and purged cross-validation, so that the validation error is honest.
- Early stopping for models fitted by iteration; see gradient descent and its pathologies.
Overfitting is fitting noise that will not repeat, and it shows as a gap between training and validation error. On financial data the noise term dominates, so flexibility is expensive and the default outcome of a search over many features is a model that fits noise.
When asked for an example, use your own. Describe a result that looked strong in sample, the two numbers that showed the gap, and what you changed. A model you overfit and then caught shows that you know how to check.
Test your knowledge
Keep reading Machine Learning for Quantitative Research
27 lessons in this course, and every other premium course, on one subscription.
- Every lesson in every course, with the worked examples and interactive simulators
- Graded questions on every lesson, with explanations for the wrong answers as well as the right one
- The trainers, timed assessments and brainteaser library that go with them