Skip to content
← The Scroll
Contemporary · Paper

Language Models are Few-Shot Learners (GPT-3)

Brown et al. (OpenAI) · 2020

"Scaling language models to 175B parameters yields emergent few-shot task performance without fine-tuning."

The idea

GPT-3's core training task never changed from earlier language models — predict the next word in a sequence, over and over, on a vast amount of internet text. What changed was scale: 175 billion parameters, trained on an enormous dataset. At that scale, something unexpected showed up: the model could perform tasks it was never specifically trained to do — translation, arithmetic, simple coding — just by being shown a few examples of the task directly inside the prompt, with no retraining or fine-tuning required.

Why it works

The mechanism demonstrated is 'in-context learning': instead of updating the model's internal weights to learn a new task (the traditional fine-tuning approach), you simply describe the task or show a handful of example input-output pairs within the prompt itself, and the model's existing attention mechanisms adapt to the pattern on the fly, producing appropriate output for a new, unseen input in the same format — without any gradient updates or retraining. This proved that scale itself could unlock general-purpose capability from a single model, rather than needing a separate specially-trained model for each task, which reframed AI in the eyes of investors from a collection of narrow tools into a general-purpose platform.

The takeaway — recall it first
Check your understanding

What is 'in-context learning,' as demonstrated by GPT-3?

Further reading

Read more about the topic

The explanation above is written with AI assistance. These are the originals — go to them to check it.

  • Language Models are Few-Shot Learners (original paper)arXiv, 2020
  • How GPT3 Works — Visualizations and AnimationsJay Alammar
Up NextSuggested: Continues the theme of AI / ML Papers

Training Compute-Optimal Large Language Models (Chinchilla)

"For a fixed compute budget, models should be smaller and trained on far more data than prior practice."

Hoffmann et al. (DeepMind) · PaperContinue→
Listen
0 / 5