Efficient Estimation of Word Representations (word2vec)
Mikolov et al. (Google) · 2013
"Words can be embedded in vector space where semantic relationships become linear arithmetic."
Before word2vec, computers mostly treated words as arbitrary, unrelated symbols — 'cat' and 'dog' had no more mathematical relationship to each other than 'cat' and 'stapler.' Mikolov and colleagues at Google showed that by training a simple neural network to predict a word from its surrounding context (or vice versa) across huge amounts of text, you get a byproduct that turns out to be far more valuable than the prediction task itself: every word ends up represented as a point (a vector) in a high-dimensional space, positioned such that words used in similar contexts end up near each other, and consistent relationships (like 'male-to-female' or 'country-to-capital') show up as consistent directions you can add or subtract.
Read more about the topic
The Bitter Lesson
"General methods that leverage computation ultimately beat human-knowledge-engineered approaches in AI."