Training Neural Networks in Practice
"Your network's loss does not go down. Walk me through how you would debug it." A complete answer follows an order: check the loss before any training, prove the model can memorise a handful of examples, and only then add data and regularisation. This lesson covers the choices behind that routine. The basics of gradient descent, momentum and early stopping are in gradient descent and its pathologies.
Initialisation: why scale matters
If the weights are too small, the activations shrink layer by layer. If they are too large, they grow. The gradients do the same on the way back, so a deep network starts frozen or exploding.
Take a ReLU layer with inputs and weights of variance . The second moment of the activations is multiplied by
The rest of this lesson is for subscribers
Unlock every lesson in Machine Learning for Quantitative Research, and every other premium course.
Subscribe to continueTest your knowledge
Keep reading Machine Learning for Quantitative Research
27 lessons in this course, and every other premium course, on one subscription.
- Every lesson in every course, with the worked examples and interactive simulators
- Graded questions on every lesson, with explanations for the wrong answers as well as the right one
- The trainers, timed assessments and brainteaser library that go with them