Classification Metrics and Calibration

"Your classifier is 90% accurate. Is it any good?" Accuracy alone cannot answer that. A complete answer asks for the base rate and names the metric that fits the decision the model feeds. It then checks whether the predicted probabilities can be taken at face value. This lesson covers each metric, what it measures, and when it misleads.

The confusion matrix

Fix a threshold, call every score above it positive, and count four outcomes: true positives (TP), false positives (FP), false negatives (FN) and true negatives (TN). Three ratios follow:

Precision=TPTP+FPRecall=TPTP+FN\begin{aligned} \text{Precision} &= \frac{TP}{TP + FP} \\ \text{Recall} &= \frac{TP}{TP + FN} \end{aligned}

and the false positive rate, FP/(FP+TN)FP / (FP + TN). Precision asks how often a flag is right. Recall asks how many of the real positives were flagged. The F1 score is their harmonic mean, 2PR/(P+R)2PR / (P + R). It is one number for a fixed threshold, and it ignores true negatives entirely.

The rest of this lesson is for subscribers

Unlock every lesson in Machine Learning for Quantitative Research, and every other premium course.

Subscribe to continue

Test your knowledge

Questions are only available to subscribers.

Keep reading Machine Learning for Quantitative Research

27 lessons in this course, and every other premium course, on one subscription.

  • Every lesson in every course, with the worked examples and interactive simulators
  • Graded questions on every lesson, with explanations for the wrong answers as well as the right one
  • The trainers, timed assessments and brainteaser library that go with them