Skip to content
← The Scroll
Contemporary · Paper

Training Compute-Optimal Large Language Models (Chinchilla)

Hoffmann et al. (DeepMind) · 2022

"For a fixed compute budget, models should be smaller and trained on far more data than prior practice."

The idea

DeepMind trained over 400 language models of varying sizes to empirically map out what actually minimizes error for a given amount of compute. Their finding directly corrected the earlier OpenAI scaling-laws guidance: rather than mostly growing parameter count, model size and training data should scale up together, roughly in equal proportion. Their 70-billion-parameter 'Chinchilla' model, trained on far more data (1.4 trillion tokens) than GPT-3's 175 billion parameters had been (roughly 300 billion tokens), ended up outperforming the much larger model.

Why it works

The mechanism is a straightforward but consequential re-derivation of the compute-optimal frontier: given a fixed training compute budget, there's a specific ratio of parameters to training tokens that minimizes loss, and the earlier scaling-laws guidance had been well off that ratio, favoring oversized, undertrained models. The practical rule of thumb that emerged — roughly 20 training tokens per parameter — reshaped how every subsequent lab budgeted a training run, and shifted the industry's bottleneck away from 'how many parameters can we afford' and toward 'how much high-quality training data can we actually acquire,' a much harder problem to solve by just buying more GPUs.

The takeaway — recall it first
Check your understanding

What did the Chinchilla paper prove was wrong with prior industry practice in training large language models?

Further reading

Read more about the topic

The explanation above is written with AI assistance. These are the originals — go to them to check it.

  • Training Compute-Optimal Large Language Models (original paper)arXiv, 2022
  • Chinchilla's Wild ImplicationsLessWrong
Up NextSuggested: Continues the theme of AI / ML Papers

Language Models are Few-Shot Learners (GPT-3)

"Scaling language models to 175B parameters yields emergent few-shot task performance without fine-tuning."

Brown et al. (OpenAI) · PaperContinue→
Listen
0 / 5