Free preview
What a Model Learns
"How would you build a model to predict this?" A full answer fixes three things before it names an algorithm: what you predict, how a mistake is scored, and which set of functions you search. The algorithm comes last, because it only searches the set you chose.
The setup
Supervised learning starts from pairs : features and a target , drawn from some distribution. A candidate model is a function chosen from a family , such as all linear functions or all trees of depth three. A loss is the cost of predicting when the truth is .
The quantity you care about is the expected risk, the average loss on new data from the same distribution:
You cannot compute it, because the distribution is unknown. You can compute the empirical risk, the average loss on the training pairs:
Training picks the in with the lowest empirical risk, possibly plus a penalty. The difference is the generalisation gap. It is positive on average, because was chosen for having a low loss on exactly these points. Its training loss is an optimistic estimate of its loss on new data. That is why every model needs data it was not fitted on, and the next lesson is about the size of that gap.
The loss decides what you predict
Each loss has a best possible prediction, and it is a different summary of the distribution of given :
| Loss | Best prediction |
|---|---|
| Squared error | The conditional mean |
| Absolute error | The conditional median |
| 0-1 loss, for classes | The most probable class |
| Log loss, for probabilities | The true probability |
The choice is therefore not a technical detail. It decides which quantity the model estimates.
A loss also has a statistical reading. Suppose the noise around a candidate is Gaussian with a fixed variance . The negative log-likelihood of one observation is then plus a constant. Minimising squared error and maximising the likelihood give the same . In the same way, absolute error is the negative log-likelihood of Laplace noise. This is why least squares and maximum likelihood agree when the noise is Gaussian.
An asset returns on 99 days out of 100 and on the other day. The median daily return is . The mean is
A model trained with absolute error learns to predict , a positive return. A strategy that buys when the prediction is positive holds an asset with a negative expected return every day.
P&L is linear in returns, so for sizing a position the mean is the quantity that matters. Squared error targets the mean, but a single large day moves its estimate a lot. Practitioners often clip extreme targets or use the Huber loss, which is squared near zero and linear in the tails. Both make the estimate more stable, and both move it away from the mean. Know which side of that trade-off you are on.
Signal and noise
Write the target as a signal plus noise, , where is the true function and has mean zero and variance . Under squared error, even the true has an expected loss of . The best possible out-of-sample is the share of the variance of that the signal explains.
For returns, that share is very small. A daily cross-sectional equity signal with an information coefficient of 0.05 is often described as good. The information coefficient (IC) is the correlation between forecast and realised return. Its is . On daily returns, a model that explains even 1% of the variance out of sample can be valuable. Three things follow.
- Training loss is mostly noise. Almost all of the variance a flexible model removes from the training loss is noise, which will not repeat.
- Evaluation needs many observations. The standard error of a sample correlation near zero is about .
- Defaults from other fields fail. Large networks and wide hyperparameter searches work when is small relative to the variance of the signal. Here they mostly fit noise.
Take one asset's daily series and a forecast whose true IC is 0.05. To tell it from zero at two standard errors, you need , so
That is the bar, not the sample you need. At the estimate is as likely to fall below 0.05 as above it, so a true IC of 0.05 clears the bar only about half the time. To clear it 80% of the time, add 0.84 standard errors:
A simulation with 20,000 runs at each size gives 49% at and 80% at .
These must be independent observations. With overlapping 20-day labels, ten years of daily data contain about 2,500 rows but only about 125 non-overlapping 20-day periods. See label overlap and sample uniqueness. A cross-sectional IC has the same problem. Stocks on the same day share market and sector moves, so the effective is far below the number of stocks times the number of days.
Four choices before an algorithm
- The target. A forward return over a stated horizon, its sign, its rank in the cross-section, or its volatility. Each is a different problem.
- The loss. Choose the loss whose best prediction is the quantity you will trade on.
- The family. The set of functions the search may return. Its size sets the trade-off between bias and variance, covered in the bias-variance decomposition.
- The data. Which observations, how many of them are independent, and which are held back for validation.
A model is the function in a chosen family with the lowest loss on the training data. The loss decides which quantity it estimates: squared error gives the mean, absolute error the median. On returns the signal is a tiny share of the variance, so training loss is mostly noise and evaluation needs many independent observations.
When asked how you would model something, state the target, the loss and the family before you name an algorithm. Then say how many independent observations you have. That number often decides the answer.
Test your knowledge
Keep reading Machine Learning for Quantitative Research
27 lessons in this course, and every other premium course, on one subscription.
- Every lesson in every course, with the worked examples and interactive simulators
- Graded questions on every lesson, with explanations for the wrong answers as well as the right one
- The trainers, timed assessments and brainteaser library that go with them