Attention Is All You Need
Vaswani et al. (Google) · 2017
"The Transformer architecture, based solely on attention, outperforms recurrence/convolution for sequence modeling."
Before Transformers, the standard way to process a sequence of words (like a sentence) was recurrently — one word at a time, in order, feeding each step's output into the next. This was slow to train (you can't easily parallelize a process that depends on the previous step finishing first) and struggled to preserve context across long distances in a sentence. Vaswani and colleagues proposed processing the entire sequence at once, using a mechanism called 'self-attention' that directly calculates how relevant every word is to every other word in the sequence, regardless of how far apart they are.
The mechanism, self-attention, computes a relevance score between every pair of tokens in the input, letting the model weigh, for each word, how much every other word should influence its representation — a pronoun like 'it' can directly attend to the noun it refers to many sentences earlier, without the information having to survive a long chain of sequential steps. Because these relevance calculations for all word-pairs can be done simultaneously rather than one step at a time, Transformers parallelize far better on modern GPU hardware than the recurrent models they replaced, which is what made training on internet-scale text datasets computationally feasible in the first place — directly enabling the large language models that followed.
What made Transformers more practical to train at massive scale than the recurrent models that came before them?
Read more about the topic
The explanation above is written with AI assistance. These are the originals — go to them to check it.
- Attention Is All You Need (original paper)arXiv, 2017
- The Illustrated TransformerJay Alammar
Software 2.0
"Neural networks are a new software paradigm where code is learned from data, not written by hand."