Deep Residual Learning for Image Recognition (ResNet)
He, Zhang, Ren, Sun (Microsoft) · 2015
"Residual connections let networks train to unprecedented depths, driving a leap in vision performance."
Before ResNet, adding more layers to a deep neural network usually helped performance up to a point — and then made things worse, even on the training data itself, which ruled out overfitting as the explanation. The real problem was that training signals (gradients) had to travel backward through every layer during training, and in very deep networks that signal would shrink to almost nothing by the time it reached the earliest layers (the 'vanishing gradient' problem), making those layers nearly impossible to train properly. He, Zhang, Ren, and Sun's fix: add 'skip connections' that let the signal bypass a block of layers entirely if needed, so a layer only has to learn the difference (the 'residual') from what came before, rather than reconstructing the whole transformation from scratch.
The mechanism is structural, not just a training trick: instead of asking each block of layers to learn a completely new mapping from input to output, ResNet reframes the task as learning a residual — the specific correction needed on top of just passing the input straight through unchanged. If a given block turns out not to be useful, the network can effectively learn to make it a no-op (its weights get pushed toward zero) and the skip connection carries the original signal straight through undamaged. This let networks scale from roughly 16 layers to over 150 layers, and the 152-layer version achieved a top-5 error rate of 3.57% on ImageNet — surpassing typical human-level performance on that specific benchmark.
What problem did ResNet's 'skip connections' specifically solve?
Read more about the topic
The explanation above is written with AI assistance. These are the originals — go to them to check it.
- Deep Residual Learning for Image Recognition (original paper)arXiv, 2015
- Understanding ResNet and its variantsTowards Data Science
Software 2.0
"Neural networks are a new software paradigm where code is learned from data, not written by hand."