Skip to content
← The Scroll
Canonical · Paper

The Anatomy of a Large-Scale Hypertextual Web Search Engine

Sergey Brin, Lawrence Page · 1998

"Ranking web pages by the link structure of the web itself (PageRank) produces dramatically better search than keyword matching."

The idea

Search engines before Google mostly ranked pages by matching keywords — how many times did the search term appear on the page. That approach was easy to game (stuff a page with the word 'concert tickets' a thousand times) and didn't capture which pages people actually trusted. Brin and Page's insight: treat every hyperlink on the web as a vote of confidence from one page to another. A page gets a higher rank not just from having many links pointing to it, but from being linked to by other pages that are themselves highly ranked — importance flows through the link graph, recursively.

Why it works

The mechanism (PageRank) models the entire web as a graph and calculates each page's score based on a recursive definition: a page's importance equals the sum of the importance of every page linking to it, divided by how many outbound links each of those pages has (so a link from a page with only one outgoing link counts for more than a link from a page linking out to a thousand other sites). This is computed iteratively across the whole graph until the scores stabilize. The key structural advantage over keyword matching is that PageRank is much harder to manipulate directly — you can't just repeat a word on your own page, you need other already-trusted pages to link to you, which is a much higher bar to fake at scale.

The takeaway — recall it first
Check your understanding

Why is PageRank harder to manipulate than simple keyword-frequency ranking?

Further reading

Read more about the topic

The explanation above is written with AI assistance. These are the originals — go to them to check it.

  • The Anatomy of a Large-Scale Hypertextual Web Search Engine (original paper)Stanford InfoLab, 1998
  • The Anatomy of a Large-Scale Hypertextual Web Search Engine (mirror)Stanford SNAP
Up NextSuggested: Continues the theme of Foundational Tech

Information Management: A Proposal

"A linked hypertext system over a network would let information at CERN (and everywhere) be shared without central control."

Tim Berners-Lee · PaperContinue→
Listen
0 / 7