Attention Is All You Need
Vaswani et al. (Google) · 2017
"The Transformer architecture, based solely on attention, outperforms recurrence/convolution for sequence modeling."
Before Transformers, the standard way to process a sequence of words (like a sentence) was recurrently — one word at a time, in order, feeding each step's output into the next. This was slow to train (you can't easily parallelize a process that depends on the previous step finishing first) and struggled to preserve context across long distances in a sentence. Vaswani and colleagues proposed processing the entire sequence at once, using a mechanism called 'self-attention' that directly calculates how relevant every word is to every other word in the sequence, regardless of how far apart they are.
Read more about the topic
Software 2.0
"Neural networks are a new software paradigm where code is learned from data, not written by hand."