Training Compute-Optimal Large Language Models (Chinchilla)
Hoffmann et al. (DeepMind) · 2022
"For a fixed compute budget, models should be smaller and trained on far more data than prior practice."
DeepMind trained over 400 language models of varying sizes to empirically map out what actually minimizes error for a given amount of compute. Their finding directly corrected the earlier OpenAI scaling-laws guidance: rather than mostly growing parameter count, model size and training data should scale up together, roughly in equal proportion. Their 70-billion-parameter 'Chinchilla' model, trained on far more data (1.4 trillion tokens) than GPT-3's 175 billion parameters had been (roughly 300 billion tokens), ended up outperforming the much larger model.
The mechanism is a straightforward but consequential re-derivation of the compute-optimal frontier: given a fixed training compute budget, there's a specific ratio of parameters to training tokens that minimizes loss, and the earlier scaling-laws guidance had been well off that ratio, favoring oversized, undertrained models. The practical rule of thumb that emerged — roughly 20 training tokens per parameter — reshaped how every subsequent lab budgeted a training run, and shifted the industry's bottleneck away from 'how many parameters can we afford' and toward 'how much high-quality training data can we actually acquire,' a much harder problem to solve by just buying more GPUs.
What did the Chinchilla paper prove was wrong with prior industry practice in training large language models?
Read more about the topic
The explanation above is written with AI assistance. These are the originals — go to them to check it.
- Training Compute-Optimal Large Language Models (original paper)arXiv, 2022
- Chinchilla's Wild ImplicationsLessWrong
Language Models are Few-Shot Learners (GPT-3)
"Scaling language models to 175B parameters yields emergent few-shot task performance without fine-tuning."