Skip to content
← Home
Contemporary · Paper

Training Compute-Optimal Large Language Models (Chinchilla)

Hoffmann et al. (DeepMind) · 2022

"For a fixed compute budget, models should be smaller and trained on far more data than prior practice."

The idea

DeepMind trained over 400 language models of varying sizes to empirically map out what actually minimizes error for a given amount of compute. Their finding directly corrected the earlier OpenAI scaling-laws guidance: rather than mostly growing parameter count, model size and training data should scale up together, roughly in equal proportion. Their 70-billion-parameter 'Chinchilla' model, trained on far more data (1.4 trillion tokens) than GPT-3's 175 billion parameters had been (roughly 300 billion tokens), ended up outperforming the much larger model.

Why it works
The takeaway — recall it first
Further reading

Read more about the topic

Up NextSuggested: Continues the theme of AI / ML Papers

Language Models are Few-Shot Learners (GPT-3)

"Scaling language models to 175B parameters yields emergent few-shot task performance without fine-tuning."

Brown et al. (OpenAI) · PaperContinue→
Listen
0 / 3