Skip to content
← The Scroll
Contemporary · Paper

Efficient Estimation of Word Representations (word2vec)

Mikolov et al. (Google) · 2013

"Words can be embedded in vector space where semantic relationships become linear arithmetic."

The idea

Before word2vec, computers mostly treated words as arbitrary, unrelated symbols — 'cat' and 'dog' had no more mathematical relationship to each other than 'cat' and 'stapler.' Mikolov and colleagues at Google showed that by training a simple neural network to predict a word from its surrounding context (or vice versa) across huge amounts of text, you get a byproduct that turns out to be far more valuable than the prediction task itself: every word ends up represented as a point (a vector) in a high-dimensional space, positioned such that words used in similar contexts end up near each other, and consistent relationships (like 'male-to-female' or 'country-to-capital') show up as consistent directions you can add or subtract.

Why it works

The mechanism comes in two architectures: Continuous Bag-of-Words (CBOW), which predicts a target word from the words around it, and Skip-gram, which does the reverse — predicting the surrounding context words from a single target word. Neither prediction task is the actual goal; the goal is the internal numerical representation (embedding) the network is forced to build in order to get good at that prediction task at all. Because words that appear in similar contexts get pushed to similar locations in this vector space, the geometry ends up encoding real semantic and relational structure — enabling vector arithmetic like king − man + woman ≈ queen — purely as an emergent property of the prediction objective, without anyone explicitly programming in the concept of gender or royalty.

The takeaway — recall it first
Check your understanding

How does word2vec end up capturing semantic relationships between words, if it's only trained to predict nearby words?

Further reading

Read more about the topic

The explanation above is written with AI assistance. These are the originals — go to them to check it.

  • Efficient Estimation of Word Representations in Vector Space (original paper)arXiv, 2013
  • The Illustrated Word2vecJay Alammar
Up NextSuggested: Continues the theme of AI / ML Papers

The Bitter Lesson

"General methods that leverage computation ultimately beat human-knowledge-engineered approaches in AI."

Rich Sutton · EssayContinue→
Listen
0 / 5