Free preview

Overfitting and the Train-Test Gap

A typical interview question is: "What is overfitting, and how would you know if your model is doing it?" A full answer has four parts. It gives the definition and the diagnosis. It explains why financial data makes overfitting so easy. And it gives an example from your own work.

The definition

A model overfits when it fits the noise in its training sample as well as the signal. It then performs worse on new data than its fit on the training data suggests. A more useful working version: a model overfits when a simpler model would do better out of sample.

Every sample contains both signal, which repeats in new data, and noise, which does not. A flexible enough model can fit both. The part of the fit that came from noise is lost the moment the model meets new data.

Diagnose it from two numbers

Compare the error on the data the model was trained on with the error on data it has not seen.

Training error Validation error Diagnosis
Low Much higher Overfitting: high variance
High Close to training Underfitting: high bias
Low Close to training A model that generalises

On returns, the same comparison appears as an in-sample Sharpe ratio of 3 that falls to 0.4 out of sample, or a training R2R^2 of 0.40 against 0.02 on validation. The gap between the two numbers is the size of the overfit.

A learning curve adds a second view: plot both errors as the training set grows. If the gap narrows as data is added, the model has high variance and more data will help. If both errors level off high and close together, the model is too simple and more data will not help.

Bias, variance and noise

For squared error, the expected error of a model's prediction on new data splits into three parts:

E[(y−f^(x))2]=Bias2+Variance+σ2\begin{aligned} E\left[(y - \hat{f}(x))^2\right] &= \text{Bias}^2 \\ &+ \text{Variance} + \sigma^2 \end{aligned}

Bias is how far the model's average prediction is from the truth. Variance is how much the prediction changes from one training sample to another. σ2\sigma^2 is the noise that no model can remove. A more flexible model lowers bias and raises variance. The next lesson, the bias-variance decomposition, derives this split and works it through by hand.

In most prediction problems outside finance, the noise term is small and flexibility pays. When predicting returns, the noise term dominates: a daily return is mostly noise, and a model with a small out-of-sample R2R^2 can be valuable. With so little signal, the variance that flexibility adds usually costs more than the bias it removes. That is the main reason simple models are hard to beat on financial data.

Many features, few observations

Overfitting is guaranteed when features are plentiful relative to observations. Fit least squares with pp features of pure noise, plus an intercept, to nn observations of a target that is also noise. The expected in-sample R2R^2 is

E[R2]=pn−1E[R^2] = \frac{p}{n - 1}
Worked example: fifty noise features

With p=50p = 50 random features and n=101n = 101 observations, the expected in-sample R2R^2 is 50/100=0.550 / 100 = 0.5. Half the variance of a target that has no relationship with any feature appears to be explained. With p=100p = 100, the regression fits the training data perfectly.

Out of sample, the same model's R2R^2 is below zero. Forecast metrics for returns defines that measure and the out-of-sample Sharpe ratio. The predictions vary, but not with the target, so they add error to simply predicting the mean.

Selection makes it worse. Suppose a research process tries 100 noise features and keeps the 50 with the largest absolute correlation with the target. The in-sample R2R^2 of the refit is then about 0.65, not 0.5. That is the average over 2,000 simulated runs with n=101n = 101. On 1,000 new observations of the same noise, the R2R^2 was below zero in every run.

The noise benchmark

The most convincing test of a whole research pipeline is to run it on data that contains no signal. Shuffle the target, or replace it with random numbers, and run the same feature selection, tuning and validation. If the pipeline still reports good out-of-sample results, something in it is leaking information or selecting on noise. The noise benchmark catches problems that no single statistic does, because it tests the process rather than one model.

What reduces overfitting

Key takeaway

Overfitting is fitting noise that will not repeat, and it shows as a gap between training and validation error. On financial data the noise term dominates, so flexibility is expensive and the default outcome of a search over many features is a model that fits noise.

Tip

When asked for an example, use your own. Describe a result that looked strong in sample, the two numbers that showed the gap, and what you changed. A model you overfit and then caught shows that you know how to check.

Test your knowledge

A model of daily returns has a training \( R^2 \) of 0.18 and a validation \( R^2 \) of 0.005. What is the most likely diagnosis?
Least squares is fitted with 30 features of pure noise, plus an intercept, to 121 observations of a target that is also pure noise. What in-sample \( R^2 \) should you expect on average?
Why do low signal-to-noise problems, such as predicting daily returns, tend to favour simpler models?

Keep reading Machine Learning for Quantitative Research

27 lessons in this course, and every other premium course, on one subscription.

  • Every lesson in every course, with the worked examples and interactive simulators
  • Graded questions on every lesson, with explanations for the wrong answers as well as the right one
  • The trainers, timed assessments and brainteaser library that go with them