Training Compute-Optimal Large Language Models (Chinchilla)
Hoffmann et al. (DeepMind) · 2022
"For a fixed compute budget, models should be smaller and trained on far more data than prior practice."
DeepMind trained over 400 language models of varying sizes to empirically map out what actually minimizes error for a given amount of compute. Their finding directly corrected the earlier OpenAI scaling-laws guidance: rather than mostly growing parameter count, model size and training data should scale up together, roughly in equal proportion. Their 70-billion-parameter 'Chinchilla' model, trained on far more data (1.4 trillion tokens) than GPT-3's 175 billion parameters had been (roughly 300 billion tokens), ended up outperforming the much larger model.
Read more about the topic
Language Models are Few-Shot Learners (GPT-3)
"Scaling language models to 175B parameters yields emergent few-shot task performance without fine-tuning."