Free preview
The Bias-Variance Decomposition
"Derive the bias-variance decomposition." Writing down the result is the easy part. A complete answer derives it, says where each assumption is used, and applies it to a real estimate. This lesson does all three.
The setup
The target is a fixed function plus noise, , with and . A model is fitted on a random training set . Fix a point and draw a new there. The noise in that new is independent of .
There are two sources of randomness: the training set, which makes random, and the noise in the new . The expected squared error averages over both.
The derivation
Write for the average prediction over training sets, and drop the to keep the lines short. Add and subtract and :
Square and take expectations. The three squared terms give the result. Each cross term is zero, for a stated reason:
- , because and is a constant.
- , because is independent of , so the expectation factors into .
- , because is a constant and by definition.
What remains is
The three terms are the irreducible noise, the squared bias and the variance. The decomposition holds at each point . The overall error averages it over the distribution of .
The result has two limits. The clean additive split is a property of squared error: for 0-1 loss there is no decomposition of this form. And the derivation assumes that new data comes from the same distribution as the training data. A drifting market adds error that none of the three terms describes.
A worked case: shrinking a mean
The same algebra applies to any estimator, and the smallest example is the most useful one in finance.
You estimate an asset's mean daily return from daily returns with standard deviation . The sample mean has no bias and variance . Consider the shrunk estimate with . Its bias is and its variance is , so
Setting the derivative to zero gives
Take , and one year of data, . Working in per cent, and , so . The best estimate shrinks the sample mean by more than 60%. Its error is , against for the sample mean: a reduction of about 62%.
The optimal factor depends on , which is the unknown. In practice the amount of shrinkage comes from a prior or from cross-validation, which is the subject of shrinkage estimators and empirical Bayes and of ridge, lasso and elastic net. The lesson of the example stands: when the noise is large next to the signal, accepting bias to cut variance lowers the error.
What capacity does to each term
As a model family grows, training error falls steadily. Bias falls too, because the family can get closer to . Variance rises, because the fit follows the particular noise in each training set. The test error is their sum plus , so it typically falls and then rises: the familiar U-shaped curve.
Two common knobs show this directly:
- Nearest neighbours. Averaging the targets of the nearest neighbours gives a variance of . A larger lowers variance and raises bias, because the neighbours are further from .
- Averaging models. Bagging averages models fitted on resampled data. It leaves bias nearly unchanged and lowers variance, which is why it helps unstable models such as deep trees. See decision trees and bagging.
Very large models can show a second descent: test error can fall again past the point where the model fits the training data exactly. This "double descent" is documented for large networks and for minimum-norm ("ridgeless") linear regression. Whether it helps on a few thousand noisy daily observations is contested. Kelly, Malamud and Zhou, in "The Virtue of Complexity in Return Prediction", argue that it does for return prediction. Treat it as a claim to test on your own data, not as a default.
Expected squared error at a point is noise plus squared bias plus variance. The cross terms vanish because the noise has mean zero and is independent of the training set, and because the prediction's deviation from its average has mean zero. When noise dominates, as on returns, a biased estimator with lower variance usually has the lower error.
At the whiteboard, define first, add and subtract, and name the reason each cross term is zero. Then give one example of the trade-off, such as shrinking a mean or choosing the number of neighbours. Where the assumptions enter matters more than the final line.
Test your knowledge
Keep reading Machine Learning for Quantitative Research
27 lessons in this course, and every other premium course, on one subscription.
- Every lesson in every course, with the worked examples and interactive simulators
- Graded questions on every lesson, with explanations for the wrong answers as well as the right one
- The trainers, timed assessments and brainteaser library that go with them