Skip to content
← Home
Contemporary · Paper

Attention Is All You Need

Vaswani et al. (Google) · 2017

"The Transformer architecture, based solely on attention, outperforms recurrence/convolution for sequence modeling."

The idea

Before Transformers, the standard way to process a sequence of words (like a sentence) was recurrently — one word at a time, in order, feeding each step's output into the next. This was slow to train (you can't easily parallelize a process that depends on the previous step finishing first) and struggled to preserve context across long distances in a sentence. Vaswani and colleagues proposed processing the entire sequence at once, using a mechanism called 'self-attention' that directly calculates how relevant every word is to every other word in the sequence, regardless of how far apart they are.

Why it works
The takeaway — recall it first
Further reading

Read more about the topic

Up NextSuggested: Continues the theme of AI / ML Papers

Software 2.0

"Neural networks are a new software paradigm where code is learned from data, not written by hand."

Andrej Karpathy · EssayContinue→
Listen
0 / 3