Language Models are Few-Shot Learners (GPT-3)
Brown et al. (OpenAI) · 2020
"Scaling language models to 175B parameters yields emergent few-shot task performance without fine-tuning."
GPT-3's core training task never changed from earlier language models — predict the next word in a sequence, over and over, on a vast amount of internet text. What changed was scale: 175 billion parameters, trained on an enormous dataset. At that scale, something unexpected showed up: the model could perform tasks it was never specifically trained to do — translation, arithmetic, simple coding — just by being shown a few examples of the task directly inside the prompt, with no retraining or fine-tuning required.
The mechanism demonstrated is 'in-context learning': instead of updating the model's internal weights to learn a new task (the traditional fine-tuning approach), you simply describe the task or show a handful of example input-output pairs within the prompt itself, and the model's existing attention mechanisms adapt to the pattern on the fly, producing appropriate output for a new, unseen input in the same format — without any gradient updates or retraining. This proved that scale itself could unlock general-purpose capability from a single model, rather than needing a separate specially-trained model for each task, which reframed AI in the eyes of investors from a collection of narrow tools into a general-purpose platform.
What is 'in-context learning,' as demonstrated by GPT-3?
Read more about the topic
The explanation above is written with AI assistance. These are the originals — go to them to check it.
- Language Models are Few-Shot Learners (original paper)arXiv, 2020
- How GPT3 Works — Visualizations and AnimationsJay Alammar
Training Compute-Optimal Large Language Models (Chinchilla)
"For a fixed compute budget, models should be smaller and trained on far more data than prior practice."