Neural Networks in Brief

Few research interviews ask for deep learning in depth, but one question comes up whenever a CV mentions it: "explain how it works, and tell me what goes wrong on financial data". The risk for the candidate is naming a method they cannot explain. This lesson covers the mechanism at the level an interviewer expects, and the pitfalls that matter most when data is noisy and scarce.

What a network computes

A feed-forward network is a chain of layers. Each layer applies a linear map, z=Wx+bz = Wx + b, and then a non-linear activation to each element, most commonly the ReLU, max⁡(0,z)\max(0, z). The output of one layer is the input to the next, and the last layer produces the prediction.

The activation is what gives the network its power. Without it, a stack of linear maps is itself one linear map, however many layers there are, and the network is just linear regression. With it, a network with one hidden layer of enough units can approximate any continuous function on a bounded range of inputs to any accuracy. That result says a network can represent the truth. It says nothing about whether it can learn it from the data available.

The rest of this lesson is for subscribers

Unlock every lesson in Advanced Topics in Probability and Statistics, and every other premium course.

Subscribe to continue

Test your knowledge

Questions are only available to subscribers.

Keep reading Advanced Topics in Probability and Statistics

47 lessons in this course, and every other premium course, on one subscription.

  • Every lesson in every course, with the worked examples and interactive simulators
  • Graded questions on every lesson, with explanations for the wrong answers as well as the right one
  • The trainers, timed assessments and brainteaser library that go with them