Recurrent networks read a sequence one step at a time, which makes them slow to train and weak at long-range dependencies. This article walks through how the field got from RNNs to the Transformer, step by step:

  • Why RNNs struggle with vanishing gradients, and how LSTMs partly fix it
  • What attention is, and why it lets a model focus on the relevant parts of the input
  • How the Transformer from “Attention Is All You Need” drops recurrence for self-attention and processes sequences in parallel
  • The encoder-decoder blocks: embeddings, positional encoding, multi-head attention and feed-forward layers

Read the full article on Medium

Updated: