Risk-Adjusted Evaluation
The final step of validation evaluates what will actually run: the sized, costed portfolio, not the classifier that generated it. Metrics inherited from machine learning (accuracy, AUC, RMSE) are useful during development and misleading as verdicts, and knowing exactly why is a reliable interview differentiator.
Why accuracy misleads
Three gaps separate classifier metrics from money.
Magnitude. Accuracy weights every observation equally; PnL weights by size of move and size of position. A model right on 55% of small moves and wrong on the large ones has good accuracy and loses money; the hit-rate-against-payoff pairing already made this point from the tear-sheet side.
Costs and constraints. A forecast is only monetisable if the trade it implies survives costs, borrow, and position limits. High-frequency reversal signals routinely show excellent AUC on moves smaller than the spread.
The rest of this lesson is for subscribers
Unlock every lesson in Research Validation and Backtesting, and every other premium course.
Subscribe to continueTest your knowledge
Keep reading Research Validation and Backtesting
21 lessons in this course, and every other premium course, on one subscription.
- Every lesson in every course, with the worked examples and interactive simulators
- Graded questions on every lesson, with explanations for the wrong answers as well as the right one
- The trainers, timed assessments and brainteaser library that go with them