Sequence Models and Attention

"Explain self-attention. Why are the scores divided by the square root of the dimension? Would you use a transformer to forecast daily returns?" A full answer derives the scaling in two lines, explains how a sequence model can see the future, and counts observations against parameters. Neural networks in brief gives the one-sentence version of each model. This lesson gives the mechanism.

Recurrent networks and the vanishing gradient

A recurrent network reads a sequence one step at a time and carries a hidden state:

ht=tanh⁡(Wht−1+Uxt+b)h_t = \tanh(W h_{t-1} + U x_t + b)

Training unrolls the network through time and backpropagates. The gradient of a loss at step tt with respect to the state τ\tau steps earlier is a product of τ\tau Jacobian matrices, one per step.

A product of many factors either shrinks or grows. If each step scales the gradient by about 0.9, then after 50 steps 0.950≈0.0050.9^{50} \approx 0.005 of it is left. At 1.1 per step it grows by 1.150≈1171.1^{50} \approx 117. A plain recurrent network learns little from events more than a few dozen steps back, or its training blows up. Gradient clipping, which rescales any gradient whose norm exceeds a threshold, controls the explosion. It does nothing for the vanishing.

The rest of this lesson is for subscribers

Unlock every lesson in Machine Learning for Quantitative Research, and every other premium course.

Subscribe to continue

Test your knowledge

Questions are only available to subscribers.

Keep reading Machine Learning for Quantitative Research

27 lessons in this course, and every other premium course, on one subscription.

  • Every lesson in every course, with the worked examples and interactive simulators
  • Graded questions on every lesson, with explanations for the wrong answers as well as the right one
  • The trainers, timed assessments and brainteaser library that go with them