Efficient Estimation of Word Representations (word2vec)
Mikolov et al. (Google) · 2013
"Words can be embedded in vector space where semantic relationships become linear arithmetic."
Before word2vec, computers mostly treated words as arbitrary, unrelated symbols — 'cat' and 'dog' had no more mathematical relationship to each other than 'cat' and 'stapler.' Mikolov and colleagues at Google showed that by training a simple neural network to predict a word from its surrounding context (or vice versa) across huge amounts of text, you get a byproduct that turns out to be far more valuable than the prediction task itself: every word ends up represented as a point (a vector) in a high-dimensional space, positioned such that words used in similar contexts end up near each other, and consistent relationships (like 'male-to-female' or 'country-to-capital') show up as consistent directions you can add or subtract.
The mechanism comes in two architectures: Continuous Bag-of-Words (CBOW), which predicts a target word from the words around it, and Skip-gram, which does the reverse — predicting the surrounding context words from a single target word. Neither prediction task is the actual goal; the goal is the internal numerical representation (embedding) the network is forced to build in order to get good at that prediction task at all. Because words that appear in similar contexts get pushed to similar locations in this vector space, the geometry ends up encoding real semantic and relational structure — enabling vector arithmetic like king − man + woman ≈ queen — purely as an emergent property of the prediction objective, without anyone explicitly programming in the concept of gender or royalty.
How does word2vec end up capturing semantic relationships between words, if it's only trained to predict nearby words?
Read more about the topic
The explanation above is written with AI assistance. These are the originals — go to them to check it.
- Efficient Estimation of Word Representations in Vector Space (original paper)arXiv, 2013
- The Illustrated Word2vecJay Alammar
The Bitter Lesson
"General methods that leverage computation ultimately beat human-knowledge-engineered approaches in AI."