Transformer Neural Network: Step-By-Step Breakdown of the Beast
Recurrent networks read a sequence one step at a time, which makes them slow to train and weak at long-range dependencies. This article walks through how the field got from RNNs to the Transformer, step by step:
- Why RNNs struggle with vanishing gradients, and how LSTMs partly fix it
- What attention is, and why it lets a model focus on the relevant parts of the input
- How the Transformer from “Attention Is All You Need” drops recurrence for self-attention and processes sequences in parallel
- The encoder-decoder blocks: embeddings, positional encoding, multi-head attention and feed-forward layers