Free preview

Most Backtests Are False Positives

Start with the honest question: before running any test, what fraction of the strategy ideas a researcher tries are genuinely profitable after costs? In liquid, well-arbitraged markets the base rate is low. Most plausible-sounding ideas are already priced in, already decayed, or were never real. That single fact drives everything in this course.

The arithmetic of low base rates

Suppose 1 in 50 candidate ideas is real, and the validation process has an 80% chance of confirming a real edge and a 5% chance of endorsing a false one. Test 1,000 ideas and validation approves 20×0.8=1620 \times 0.8 = 16 real strategies and 980×0.05=49980 \times 0.05 = 49 false ones.

P(realapproved)=1616+4925%P(\text{real} \mid \text{approved}) = \frac{16}{16 + 49} \approx 25\%

Three out of four approved strategies are false positives, and nothing went wrong: no leakage, no bugs, no bad faith. A 5% false-approval rate feels rigorous and still loses to the base rate. This is the research version of the classic diagnostic-test problem, and interviewers use it in exactly that form.

Key takeaway

"Statistically significant" answers the wrong question. Significance controls how often noise passes the test; what you care about is the probability the strategy is real given that it passed, and that depends on the base rate and on how many things were tried.

Selection makes it worse

The arithmetic above assumes every trial is reported. Research does the opposite: the losers are quietly discarded and the single best result is polished and presented. Selecting the best of many attempts inflates the winner mechanically, an effect this course quantifies later through the deflated Sharpe ratio.

Worked example: the drawer of dead strategies

A researcher spends a quarter testing 60 variations of a momentum idea and presents the one with an in-sample Sharpe of 1.3. If the 60 results were mean-zero noise with standard deviation 0.5 in Sharpe units, the expected best of 60 sits near 0.5×2.31.20.5 \times 2.3 \approx 1.2. The presented 1.3 is barely distinguishable from picking the luckiest of five dozen coin flips, and the 59 dead strategies in the drawer are the missing context.

What follows for practice

Two habits fall directly out of the arithmetic. First, an economic prior is part of the evidence: an idea with a defensible reason to exist (a risk premium, a structural flow, a behavioural constraint) starts from a higher base rate than a pattern found by search. Second, the number of trials behind a result is a first-class statistic, worth recording and reporting alongside the Sharpe.

The rest of the course dismantles the failure modes one family at a time. The next lesson names the taxonomy the following sections build on.

Test your knowledge

Suppose 1 in 50 candidate strategy ideas is genuinely profitable, validation confirms a real edge 80% of the time, and it wrongly endorses a false one 5% of the time. Of 1,000 ideas tested, roughly what fraction of the approved strategies are real?
Why is "the result is statistically significant at 5%" not an answer to "is this strategy real?"
Why does an economic rationale (a risk premium, a structural flow, a behavioural constraint) strengthen a backtest result?

Keep reading Research Validation and Backtesting

21 lessons in this course, and every other premium course, on one subscription.

  • Every lesson in every course, with the worked examples and interactive simulators
  • Graded questions on every lesson, with explanations for the wrong answers as well as the right one
  • The trainers, timed assessments and brainteaser library that go with them