Residual networks
A residual block adds a learned residual function F(x) to a shortcut carrying the block input x.
The original mapping is recast as F(x) + x. The stacked layers learn the residual mapping F(x). A shortcut connection and element-wise addition perform the F(x) + x operation.
Residual learning addressed the degradation problem observed in deeper plain networks. Identity shortcut connections add neither extra parameters nor computational complexity. When dimensions differ, the shortcut can use a linear projection.
Wide residual networks multiply the number of convolutional features by a factor k. ResNeXt calls the size of the set of transformations cardinality. Fixup enables very deep residual networks to train without normalization.
Adding layers to a suitably deep plain network can increase its training error.
Follow one activation through a residual block.
- 1 · shortcutA shortcut connection skips one or more layers.
- 2 · transformThe stacked layers in the residual branch compute F(x).
- 3 · carryIdentity shortcut connections add neither extra parameters nor computational complexity.
- 4 · addElement-wise addition combines F(x) and x.
If the dimensions differ, the shortcut can use a learned linear projection before addition.
| Who | What they ask | What it works with |
|---|---|---|
| Vision researcher | “Can a deeper image classifier avoid the plain network's degradation problem?” | Residual blocks in a convolutional network |
| Architecture researcher | “What happens if each residual stage is widened?” | Feature width multiplier |
| Optimisation researcher | “Can a residual network train without normalization?” | Fixup initialization rules |
- Residual learning addressed the degradation problem observed in deeper plain networks.
- Identity shortcuts add no extra parameters or computational complexity.
- Identity mappings provide a direct route for forward and backward signals between blocks.
- The tensors F(x) and x must have equal dimensions for direct element-wise addition.
- The original formulation uses a projection shortcut when dimensions change.
- In Wide ResNets, parameter count and computational complexity are quadratic in the width multiplier.
- Fixup was introduced because standard initialization without normalization can cause exploding gradients in residual networks.
Sources used
This explainer is written in original language. The links below support its factual claims.
- paperDeep Residual Learning for Image Recognition, He, Zhang, Ren and Sun · read 28 Sept 2026
- paperIdentity Mappings in Deep Residual Networks, He et al. · read 28 Sept 2026
- paperWide Residual Networks, Zagoruyko and Komodakis · read 28 Sept 2026
- paperAggregated Residual Transformations for Deep Neural Networks, Xie et al. · read 28 Sept 2026
- paperFixup Initialization, Zhang, Dauphin and Ma · read 28 Sept 2026