Scaling Laws for Neural Language Models
Kaplan et al. (OpenAI) · 2020
"Model performance scales predictably as a power law with compute, data, and parameters."
Kaplan and colleagues at OpenAI ran systematic experiments training many language models of different sizes and found something remarkably clean: a model's performance (its error rate) improves in a predictable mathematical pattern (a power law) as you increase compute, dataset size, or parameter count — and crucially, the internal architecture details (like how deep versus how wide the network is) mattered far less than sheer scale. This gave labs a rare thing in machine learning research: a way to forecast how much better a bigger model would be, before actually building it.
Read more about the topic
Training Compute-Optimal Large Language Models (Chinchilla)
"For a fixed compute budget, models should be smaller and trained on far more data than prior practice."