Training Neural Networks in Practice

"Your network's loss does not go down. Walk me through how you would debug it." A complete answer follows an order: check the loss before any training, prove the model can memorise a handful of examples, and only then add data and regularisation. This lesson covers the choices behind that routine. The basics of gradient descent, momentum and early stopping are in gradient descent and its pathologies.

Initialisation: why scale matters

If the weights are too small, the activations shrink layer by layer. If they are too large, they grow. The gradients do the same on the way back, so a deep network starts frozen or exploding.

Take a ReLU layer with nn inputs and weights of variance Var(w)\text{Var}(w). The second moment of the activations is multiplied by

n Var(w)2\frac{n \, \text{Var}(w)}{2}

The rest of this lesson is for subscribers

Unlock every lesson in Machine Learning for Quantitative Research, and every other premium course.

Subscribe to continue

Test your knowledge

Questions are only available to subscribers.

Keep reading Machine Learning for Quantitative Research

27 lessons in this course, and every other premium course, on one subscription.

  • Every lesson in every course, with the worked examples and interactive simulators
  • Graded questions on every lesson, with explanations for the wrong answers as well as the right one
  • The trainers, timed assessments and brainteaser library that go with them