ImageNet Classification with Deep Convolutional Neural Networks (AlexNet)
Krizhevsky, Sutskever, Hinton · 2012
"Deep CNNs trained on GPUs shatter the ImageNet benchmark, igniting the modern deep-learning era."
Every year, the ImageNet competition tested how well computer programs could correctly label photos into categories. For years, progress came from hand-engineered techniques designed by computer vision experts. In 2012, a deep convolutional neural network trained by Krizhevsky, Sutskever, and Hinton entered and won by such a wide margin that it upended the field's consensus: a network with enough layers, trained on enough labeled images, using specific practical tricks to make training feasible, could learn better visual features on its own than experts could design by hand.
The mechanism combined several practical breakthroughs that made deep networks trainable at this scale for the first time: the ReLU activation function, which trains many times faster than the previously standard sigmoid function because it doesn't saturate (stop learning) as easily; dropout, a technique that randomly disables a fraction of neurons during training so the network can't over-rely on any single pathway, which reduces overfitting; and splitting the 60 million parameters of the model across two GPUs running in parallel, since no single GPU at the time had enough memory to hold the whole network. None of these tricks were individually brand new, but combining them let a deep network finally be trained in a practical amount of time at a scale that outperformed hand-engineered approaches.
What combination of techniques allowed AlexNet to train a deep network fast enough to be practical in 2012?
Read more about the topic
The explanation above is written with AI assistance. These are the originals — go to them to check it.
- ImageNet Classification with Deep Convolutional Neural Networks (original paper)NeurIPS, 2012
- Geoffrey Hinton — Turing Award Lecture on Deep LearningACM / YouTube
The Bitter Lesson
"General methods that leverage computation ultimately beat human-knowledge-engineered approaches in AI."