Skip to content
← The Scroll
Contemporary · Paper

Scaling Laws for Neural Language Models

Kaplan et al. (OpenAI) · 2020

"Model performance scales predictably as a power law with compute, data, and parameters."

The idea

Kaplan and colleagues at OpenAI ran systematic experiments training many language models of different sizes and found something remarkably clean: a model's performance (its error rate) improves in a predictable mathematical pattern (a power law) as you increase compute, dataset size, or parameter count — and crucially, the internal architecture details (like how deep versus how wide the network is) mattered far less than sheer scale. This gave labs a rare thing in machine learning research: a way to forecast how much better a bigger model would be, before actually building it.

Why it works

The mechanism is an empirical, not theoretical, finding: by training a large number of models across a wide range of sizes and measuring how loss (prediction error) changed, the researchers found the relationship followed a smooth power-law curve rather than being unpredictable or plateauing. Their specific recommendation — that given a fixed compute budget, you should mostly grow the model's parameter count rather than proportionally growing the training dataset — became the guiding assumption behind years of ever-larger model releases. This is also a useful cautionary case: that specific parameter-versus-data recommendation was later shown to be significantly off by DeepMind's 2022 Chinchilla paper, which found models were being kept far too large relative to how much data they were trained on.

The takeaway — recall it first
Check your understanding

What specific guidance from the original Scaling Laws paper was later overturned by DeepMind's Chinchilla paper?

Further reading

Read more about the topic

The explanation above is written with AI assistance. These are the originals — go to them to check it.

  • Scaling Laws for Neural Language Models (original paper)arXiv, 2020
  • The Scaling HypothesisGwern Branwen
Up NextSuggested: Continues the theme of AI / ML Papers

Training Compute-Optimal Large Language Models (Chinchilla)

"For a fixed compute budget, models should be smaller and trained on far more data than prior practice."

Hoffmann et al. (DeepMind) · PaperContinue→
Listen
0 / 5