Language Models are Few-Shot Learners (GPT-3)
Brown et al. (OpenAI) · 2020
"Scaling language models to 175B parameters yields emergent few-shot task performance without fine-tuning."
GPT-3's core training task never changed from earlier language models — predict the next word in a sequence, over and over, on a vast amount of internet text. What changed was scale: 175 billion parameters, trained on an enormous dataset. At that scale, something unexpected showed up: the model could perform tasks it was never specifically trained to do — translation, arithmetic, simple coding — just by being shown a few examples of the task directly inside the prompt, with no retraining or fine-tuning required.
Read more about the topic
Training Compute-Optimal Large Language Models (Chinchilla)
"For a fixed compute budget, models should be smaller and trained on far more data than prior practice."