Skip to content
← Home
Contemporary · Paper

Deep Residual Learning for Image Recognition (ResNet)

He, Zhang, Ren, Sun (Microsoft) · 2015

"Residual connections let networks train to unprecedented depths, driving a leap in vision performance."

The idea

Before ResNet, adding more layers to a deep neural network usually helped performance up to a point — and then made things worse, even on the training data itself, which ruled out overfitting as the explanation. The real problem was that training signals (gradients) had to travel backward through every layer during training, and in very deep networks that signal would shrink to almost nothing by the time it reached the earliest layers (the 'vanishing gradient' problem), making those layers nearly impossible to train properly. He, Zhang, Ren, and Sun's fix: add 'skip connections' that let the signal bypass a block of layers entirely if needed, so a layer only has to learn the difference (the 'residual') from what came before, rather than reconstructing the whole transformation from scratch.

Why it works
The takeaway — recall it first
Further reading

Read more about the topic

Up NextSuggested: Continues the theme of AI / ML Papers

Software 2.0

"Neural networks are a new software paradigm where code is learned from data, not written by hand."

Andrej Karpathy · EssayContinue→
Listen
0 / 3