Support Vector Machines and Kernels

A typical interview question is: "How does a support vector machine differ from logistic regression?" The short answer is the loss. Both fit a linear score f(x)=w⊤x+bf(x) = w^\top x + b and classify by its sign. They differ in how they score a mistake, and that difference decides which points shape the fit.

Losses as functions of the margin

Code the labels as y∈{−1,+1}y \in \{-1, +1\}. The margin of a point is m=y f(x)m = y \, f(x). It is positive when the point is on the correct side, and large when it is on the correct side by a wide distance. Both losses are functions of mm alone:

hinge(m)=max⁡(0, 1−m)logistic(m)=log⁡(1+e−m)\begin{aligned} \text{hinge}(m) &= \max(0, \, 1 - m) \\ \text{logistic}(m) &= \log(1 + e^{-m}) \end{aligned}

The logistic loss is the log loss of logistic regression, rewritten for labels of ±1\pm 1.

The rest of this lesson is for subscribers

Unlock every lesson in Machine Learning for Quantitative Research, and every other premium course.

Subscribe to continue

Test your knowledge

Questions are only available to subscribers.

Keep reading Machine Learning for Quantitative Research

27 lessons in this course, and every other premium course, on one subscription.

  • Every lesson in every course, with the worked examples and interactive simulators
  • Graded questions on every lesson, with explanations for the wrong answers as well as the right one
  • The trainers, timed assessments and brainteaser library that go with them