The Scaling Hypothesis
Gwern Branwen · 2020
"Intelligence may emerge from scaling up compute and data. GPT-3 shows the hypothesis holds."
Gwern Branwen's essay explaining why Deep Learning took over AI. The 'Scaling Hypothesis' posits that simply throwing exponentially more compute and data at simple neural network architectures (like Transformers) will inevitably lead to Artificial General Intelligence (AGI).
For decades, AI researchers believed AGI would require complex, hand-coded architectures and entirely new breakthroughs in cognitive science. Gwern points out that the 'Bitter Lesson' of AI is that simple algorithms combined with massive scale always win. The Scaling Hypothesis states that the emergent reasoning capabilities we see in large language models are not a trick; they are the fundamental result of scale. If the hypothesis holds, the path to AGI isn't a mysterious scientific breakthrough, it's just an engineering and capital allocation problem of building larger data centers.
What does the 'Scaling Hypothesis' argue is the path to AGI?
Read more about the topic
The explanation above is written with AI assistance. These are the originals — go to them to check it.
- The Scaling Hypothesis (Gwern)House overview / Gwern
The Next Big Thing Will Start Out Looking Like a Toy
"Disruptive innovations are dismissed as toys because they underperform on established metrics while excelling on new ones."