Label Overlap and Sample Uniqueness

The subtlest leakage channel is not in the features at all. It is in the labels, and it is the reason a model can validate beautifully on data it has, in effect, already trained on.

Overlapping labels

Financial labels are usually forward-looking windows: "the return over the next 20 days," "did the price hit the barrier within a week." Build one such label per day and consecutive labels overlap: today's 20-day label and tomorrow's share 19 days of the same price path. They are not two pieces of evidence; they are one piece of evidence counted almost twice.

Two consequences follow, and interviews test both.

Inflated confidence. Every standard error, t-statistic and p-value assumes independent observations. With 20-day labels sampled daily, a decade of data looks like roughly 2,500 observations but carries something closer to 125 independent ones. The nominal t-statistic on a fitted coefficient is inflated by a factor of about 20\sqrt{20}, so effects that are pure noise test as highly significant.

The rest of this lesson is for subscribers

Unlock every lesson in Research Validation and Backtesting, and every other premium course.

Subscribe to continue

Test your knowledge

Questions are only available to subscribers.

Keep reading Research Validation and Backtesting

21 lessons in this course, and every other premium course, on one subscription.

  • Every lesson in every course, with the worked examples and interactive simulators
  • Graded questions on every lesson, with explanations for the wrong answers as well as the right one
  • The trainers, timed assessments and brainteaser library that go with them