Residual connections
A residual connection combines a residual mapping F(x) with the block input x as F(x) + x.
The learned route transforms x into F(x). The block then adds the two tensors and returns F(x) + x.
An identity shortcut adds neither extra parameters nor computational complexity. Direct identity-shortcut addition works only when x and F(x) have matching dimensions. If the dimensions differ, the shortcut can use a linear projection to match the dimensions.
Transformers employ a residual connection around each of two sublayers. Highway networks regulate information flow with learned gates. Fixup enables residual networks without normalization.
Deeper neural networks can be more difficult to train.
Follow one tensor through the shortcut and learned branch.
- 1 · splitThe shortcut connection adds x to the residual mapping F(x).
- 2 · transformStacked layers compute a residual mapping F(x).
- 3 · carryDirect identity-shortcut addition requires x and F(x) to have equal dimensions.
- 4 · addElement-wise addition forms y = F(x) + x.
When dimensions differ, the shortcut can apply a linear projection before addition.
| Who | What they ask | What it works with |
|---|---|---|
| Vision engineer | “How can this block refine a feature map without discarding it?” | An identity shortcut around convolutional layers |
| Language-model engineer | “Where does each Transformer sublayer return its input?” | Residual connections around Transformer sublayers |
| Optimisation researcher | “Can a deep residual network train without normalization?” | Fixup initialization |
- Residual learning was introduced to ease the optimization of substantially deeper networks.
- An identity shortcut adds neither parameters nor computational complexity.
- Identity mappings allow forward and backward signals to pass directly between residual blocks.
- Direct identity-shortcut addition requires x and F(x) to have equal dimensions.
Sources used
This explainer is written in original language. The links below support its factual claims.
- paperDeep Residual Learning for Image Recognition, He, Zhang, Ren and Sun · read 28 Sept 2026
- paperIdentity Mappings in Deep Residual Networks, He et al. · read 28 Sept 2026
- paperHighway Networks, Srivastava, Greff and Schmidhuber · read 28 Sept 2026
- paperAttention Is All You Need, Vaswani et al. · read 28 Sept 2026
- paperFixup Initialization: Residual Learning Without Normalization, Zhang, Dauphin and Ma · read 28 Sept 2026