Free preview

What a Model Learns

"How would you build a model to predict this?" A full answer fixes three things before it names an algorithm: what you predict, how a mistake is scored, and which set of functions you search. The algorithm comes last, because it only searches the set you chose.

The setup

Supervised learning starts from pairs (x,y)(x, y): features xx and a target yy, drawn from some distribution. A candidate model is a function gg chosen from a family F\mathcal{F}, such as all linear functions or all trees of depth three. A loss L(y,y^)L(y, \hat{y}) is the cost of predicting y^\hat{y} when the truth is yy.

The quantity you care about is the expected risk, the average loss on new data from the same distribution:

R(g)=E[L(y,g(x))]R(g) = E\left[L(y, g(x))\right]

You cannot compute it, because the distribution is unknown. You can compute the empirical risk, the average loss on the nn training pairs:

R^(g)=1n∑i=1nL(yi,g(xi))\hat{R}(g) = \frac{1}{n} \sum_{i=1}^{n} L(y_i, g(x_i))

Training picks the f^\hat{f} in F\mathcal{F} with the lowest empirical risk, possibly plus a penalty. The difference R(f^)−R^(f^)R(\hat{f}) - \hat{R}(\hat{f}) is the generalisation gap. It is positive on average, because f^\hat{f} was chosen for having a low loss on exactly these nn points. Its training loss is an optimistic estimate of its loss on new data. That is why every model needs data it was not fitted on, and the next lesson is about the size of that gap.

The loss decides what you predict

Each loss has a best possible prediction, and it is a different summary of the distribution of yy given xx:

Loss Best prediction
Squared error (y−y^)2(y - \hat{y})^2 The conditional mean E[y∣x]E[y \mid x]
Absolute error ∣y−y^∣\lvert y - \hat{y} \rvert The conditional median
0-1 loss, for classes The most probable class
Log loss, for probabilities The true probability

The choice is therefore not a technical detail. It decides which quantity the model estimates.

A loss also has a statistical reading. Suppose the noise around a candidate gg is Gaussian with a fixed variance σ2\sigma^2. The negative log-likelihood of one observation is then (y−g(x))2/(2σ2)(y - g(x))^2 / (2\sigma^2) plus a constant. Minimising squared error and maximising the likelihood give the same f^\hat{f}. In the same way, absolute error is the negative log-likelihood of Laplace noise. This is why least squares and maximum likelihood agree when the noise is Gaussian.

Worked example: mean against median on skewed returns

An asset returns +0.1%+0.1\% on 99 days out of 100 and −15%-15\% on the other day. The median daily return is +0.1%+0.1\%. The mean is

99×0.1%−15%100=−0.051%\frac{99 \times 0.1\% - 15\%}{100} = -0.051\%

A model trained with absolute error learns to predict +0.1%+0.1\%, a positive return. A strategy that buys when the prediction is positive holds an asset with a negative expected return every day.

P&L is linear in returns, so for sizing a position the mean is the quantity that matters. Squared error targets the mean, but a single large day moves its estimate a lot. Practitioners often clip extreme targets or use the Huber loss, which is squared near zero and linear in the tails. Both make the estimate more stable, and both move it away from the mean. Know which side of that trade-off you are on.

Signal and noise

Write the target as a signal plus noise, y=f(x)+εy = f(x) + \varepsilon, where ff is the true function and ε\varepsilon has mean zero and variance σ2\sigma^2. Under squared error, even the true ff has an expected loss of σ2\sigma^2. The best possible out-of-sample R2R^2 is the share of the variance of yy that the signal explains.

For returns, that share is very small. A daily cross-sectional equity signal with an information coefficient of 0.05 is often described as good. The information coefficient (IC) is the correlation between forecast and realised return. Its R2R^2 is 0.052=0.00250.05^2 = 0.0025. On daily returns, a model that explains even 1% of the variance out of sample can be valuable. Three things follow.

  • Training loss is mostly noise. Almost all of the variance a flexible model removes from the training loss is noise, which will not repeat.
  • Evaluation needs many observations. The standard error of a sample correlation near zero is about 1/n1 / \sqrt{n}.
  • Defaults from other fields fail. Large networks and wide hyperparameter searches work when σ2\sigma^2 is small relative to the variance of the signal. Here they mostly fit noise.
Worked example: how many observations to see an edge

Take one asset's daily series and a forecast whose true IC is 0.05. To tell it from zero at two standard errors, you need 2/n≤0.052 / \sqrt{n} \le 0.05, so

n≥(20.05)2=1,600n \ge \left(\frac{2}{0.05}\right)^2 = 1{,}600

That is the bar, not the sample you need. At n=1,600n = 1{,}600 the estimate is as likely to fall below 0.05 as above it, so a true IC of 0.05 clears the bar only about half the time. To clear it 80% of the time, add 0.84 standard errors:

n≈(2+0.840.05)2≈3,200n \approx \left(\frac{2 + 0.84}{0.05}\right)^2 \approx 3{,}200

A simulation with 20,000 runs at each size gives 49% at n=1,600n = 1{,}600 and 80% at n=3,226n = 3{,}226.

These must be independent observations. With overlapping 20-day labels, ten years of daily data contain about 2,500 rows but only about 125 non-overlapping 20-day periods. See label overlap and sample uniqueness. A cross-sectional IC has the same problem. Stocks on the same day share market and sector moves, so the effective nn is far below the number of stocks times the number of days.

Four choices before an algorithm

  1. The target. A forward return over a stated horizon, its sign, its rank in the cross-section, or its volatility. Each is a different problem.
  2. The loss. Choose the loss whose best prediction is the quantity you will trade on.
  3. The family. The set of functions the search may return. Its size sets the trade-off between bias and variance, covered in the bias-variance decomposition.
  4. The data. Which observations, how many of them are independent, and which are held back for validation.
Key takeaway

A model is the function in a chosen family with the lowest loss on the training data. The loss decides which quantity it estimates: squared error gives the mean, absolute error the median. On returns the signal is a tiny share of the variance, so training loss is mostly noise and evaluation needs many independent observations.

Tip

When asked how you would model something, state the target, the loss and the family before you name an algorithm. Then say how many independent observations you have. That number often decides the answer.

Test your knowledge

A model is trained to minimise absolute error \( \lvert y - \hat{f}(x) \rvert \). Which summary of \( y \) given \( x \) does its best possible prediction estimate?
About how many independent observations do you need to tell an information coefficient of 0.04 from zero at two standard errors?
Why is the training loss of a fitted model, on average, lower than its loss on new data from the same distribution?

Keep reading Machine Learning for Quantitative Research

27 lessons in this course, and every other premium course, on one subscription.

  • Every lesson in every course, with the worked examples and interactive simulators
  • Graded questions on every lesson, with explanations for the wrong answers as well as the right one
  • The trainers, timed assessments and brainteaser library that go with them