Label Overlap and Sample Uniqueness
The subtlest leakage channel is not in the features at all. It is in the labels, and it is the reason a model can validate beautifully on data it has, in effect, already trained on.
Overlapping labels
Financial labels are usually forward-looking windows: "the return over the next 20 days," "did the price hit the barrier within a week." Build one such label per day and consecutive labels overlap: today's 20-day label and tomorrow's share 19 days of the same price path. They are not two pieces of evidence; they are one piece of evidence counted almost twice.
Two consequences follow, and interviews test both.
Inflated confidence. Every standard error, t-statistic and p-value assumes independent observations. With 20-day labels sampled daily, a decade of data looks like roughly 2,500 observations but carries something closer to 125 independent ones. The nominal t-statistic on a fitted coefficient is inflated by a factor of about , so effects that are pure noise test as highly significant.
The rest of this lesson is for subscribers
Unlock every lesson in Research Validation and Backtesting, and every other premium course.
Subscribe to continueTest your knowledge
Keep reading Research Validation and Backtesting
21 lessons in this course, and every other premium course, on one subscription.
- Every lesson in every course, with the worked examples and interactive simulators
- Graded questions on every lesson, with explanations for the wrong answers as well as the right one
- The trainers, timed assessments and brainteaser library that go with them