Gradient Descent and Its Pathologies

A typical interview question is: "Describe gradient descent, derive the update for linear regression, and tell me what goes wrong." Least squares has a closed form, so the question is about the method. Gradient descent takes over when the closed form is too expensive, or when there is none, as for logistic regression and every neural network.

The update

To minimise a loss L(β)L(\beta), start from a guess and repeatedly step against the gradient:

Gradient descent
β←β−η ∇L(β)\beta \leftarrow \beta - \eta \, \nabla L(\beta)

Step downhill by an amount proportional to the slope. The learning rate eta decides whether that step converges, crawls or diverges.

For least squares with nn observations, the loss and its gradient are

L(β)=1n∥y−Xβ∥2∇L(β)=−2nX⊤(y−Xβ)\begin{aligned} L(\beta) &= \frac{1}{n} \lVert y - X\beta \rVert^2 \\ \nabla L(\beta) &= -\frac{2}{n} X^\top (y - X\beta) \end{aligned}

The rest of this lesson is for subscribers

Unlock every lesson in Machine Learning for Quantitative Research, and every other premium course.

Subscribe to continue

Test your knowledge

Questions are only available to subscribers.

Keep reading Machine Learning for Quantitative Research

27 lessons in this course, and every other premium course, on one subscription.

  • Every lesson in every course, with the worked examples and interactive simulators
  • Graded questions on every lesson, with explanations for the wrong answers as well as the right one
  • The trainers, timed assessments and brainteaser library that go with them