Free preview

The Bias-Variance Decomposition

"Derive the bias-variance decomposition." Writing down the result is the easy part. A complete answer derives it, says where each assumption is used, and applies it to a real estimate. This lesson does all three.

The setup

The target is a fixed function plus noise, y=f(x)+εy = f(x) + \varepsilon, with E[ε]=0E[\varepsilon] = 0 and Var(ε)=σ2\text{Var}(\varepsilon) = \sigma^2. A model f^\hat{f} is fitted on a random training set DD. Fix a point xx and draw a new yy there. The noise in that new yy is independent of DD.

There are two sources of randomness: the training set, which makes f^(x)\hat{f}(x) random, and the noise in the new yy. The expected squared error averages over both.

The derivation

Write fˉ(x)=ED[f^(x)]\bar{f}(x) = E_D[\hat{f}(x)] for the average prediction over training sets, and drop the xx to keep the lines short. Add and subtract ff and fˉ\bar{f}:

y−f^=(y−f)⏟ε+(f−fˉ)⏟constant+(fˉ−f^)⏟depends on D\begin{aligned} y - \hat{f} &= \underbrace{(y - f)}_{\varepsilon} + \underbrace{(f - \bar{f})}_{\text{constant}} \\ &\quad + \underbrace{(\bar{f} - \hat{f})}_{\text{depends on } D} \end{aligned}

Square and take expectations. The three squared terms give the result. Each cross term is zero, for a stated reason:

  • E[ε (f−fˉ)]=0E[\varepsilon \,(f - \bar{f})] = 0, because E[ε]=0E[\varepsilon] = 0 and f−fˉf - \bar{f} is a constant.
  • E[ε (fˉ−f^)]=0E[\varepsilon \,(\bar{f} - \hat{f})] = 0, because ε\varepsilon is independent of DD, so the expectation factors into E[ε] E[fˉ−f^]=0E[\varepsilon] \, E[\bar{f} - \hat{f}] = 0.
  • E[(f−fˉ)(fˉ−f^)]=0E[(f - \bar{f})(\bar{f} - \hat{f})] = 0, because f−fˉf - \bar{f} is a constant and ED[f^]=fˉE_D[\hat{f}] = \bar{f} by definition.

What remains is

E[(y−f^)2]=σ2+(f−fˉ)2+E[(f^−fˉ)2]\begin{aligned} E\left[(y - \hat{f})^2\right] &= \sigma^2 + (f - \bar{f})^2 \\ &\quad + E\left[(\hat{f} - \bar{f})^2\right] \end{aligned}

The three terms are the irreducible noise, the squared bias and the variance. The decomposition holds at each point xx. The overall error averages it over the distribution of xx.

The result has two limits. The clean additive split is a property of squared error: for 0-1 loss there is no decomposition of this form. And the derivation assumes that new data comes from the same distribution as the training data. A drifting market adds error that none of the three terms describes.

A worked case: shrinking a mean

The same algebra applies to any estimator, and the smallest example is the most useful one in finance.

Worked example: how much to shrink one year of mean returns

You estimate an asset's mean daily return μ\mu from nn daily returns with standard deviation σ\sigma. The sample mean rˉ\bar{r} has no bias and variance σ2/n\sigma^2 / n. Consider the shrunk estimate c rˉc \, \bar{r} with 0≤c≤10 \le c \le 1. Its bias is (c−1)μ(c - 1)\mu and its variance is c2σ2/nc^2 \sigma^2 / n, so

MSE(c)=(1−c)2μ2+c2σ2n\text{MSE}(c) = (1 - c)^2 \mu^2 + c^2 \frac{\sigma^2}{n}

Setting the derivative to zero gives

c∗=μ2μ2+σ2/nc^* = \frac{\mu^2}{\mu^2 + \sigma^2 / n}

Take μ=0.05%\mu = 0.05\%, σ=1%\sigma = 1\% and one year of data, n=250n = 250. Working in per cent, μ2=0.0025\mu^2 = 0.0025 and σ2/n=0.004\sigma^2 / n = 0.004, so c∗=0.0025/0.0065≈0.38c^* = 0.0025 / 0.0065 \approx 0.38. The best estimate shrinks the sample mean by more than 60%. Its error is 0.0025×0.004/0.0065≈0.00150.0025 \times 0.004 / 0.0065 \approx 0.0015, against 0.0040.004 for the sample mean: a reduction of about 62%.

The optimal factor c∗c^* depends on μ\mu, which is the unknown. In practice the amount of shrinkage comes from a prior or from cross-validation, which is the subject of shrinkage estimators and empirical Bayes and of ridge, lasso and elastic net. The lesson of the example stands: when the noise is large next to the signal, accepting bias to cut variance lowers the error.

What capacity does to each term

As a model family grows, training error falls steadily. Bias falls too, because the family can get closer to ff. Variance rises, because the fit follows the particular noise in each training set. The test error is their sum plus σ2\sigma^2, so it typically falls and then rises: the familiar U-shaped curve.

Two common knobs show this directly:

  • Nearest neighbours. Averaging the targets of the mm nearest neighbours gives a variance of σ2/m\sigma^2 / m. A larger mm lowers variance and raises bias, because the neighbours are further from xx.
  • Averaging models. Bagging averages models fitted on resampled data. It leaves bias nearly unchanged and lowers variance, which is why it helps unstable models such as deep trees. See decision trees and bagging.

Very large models can show a second descent: test error can fall again past the point where the model fits the training data exactly. This "double descent" is documented for large networks and for minimum-norm ("ridgeless") linear regression. Whether it helps on a few thousand noisy daily observations is contested. Kelly, Malamud and Zhou, in "The Virtue of Complexity in Return Prediction", argue that it does for return prediction. Treat it as a claim to test on your own data, not as a default.

Key takeaway

Expected squared error at a point is noise plus squared bias plus variance. The cross terms vanish because the noise has mean zero and is independent of the training set, and because the prediction's deviation from its average has mean zero. When noise dominates, as on returns, a biased estimator with lower variance usually has the lower error.

Tip

At the whiteboard, define fˉ\bar{f} first, add and subtract, and name the reason each cross term is zero. Then give one example of the trade-off, such as shrinking a mean or choosing the number of neighbours. Where the assumptions enter matters more than the final line.

Test your knowledge

In the derivation, why is \( E[\varepsilon \, (\bar{f}(x) - \hat{f}(x))] = 0 \)?
A mean daily return of \( 0.03\% \) is estimated from 500 daily returns with standard deviation \( 1\% \). Which multiple \( c \) of the sample mean minimises the mean squared error?
Bagging averages many deep trees fitted on bootstrap samples. Which term of the decomposition does it mainly reduce?

Keep reading Machine Learning for Quantitative Research

27 lessons in this course, and every other premium course, on one subscription.

  • Every lesson in every course, with the worked examples and interactive simulators
  • Graded questions on every lesson, with explanations for the wrong answers as well as the right one
  • The trainers, timed assessments and brainteaser library that go with them