Concepts

Residual connections

4 min readintermediateUpdated 28 Sept 2026
1 · In one line

A residual connection combines a residual mapping F(x) with the block input x as F(x) + x.

1 · What it is

The learned route transforms x into F(x). The block then adds the two tensors and returns F(x) + x.

An identity shortcut adds neither extra parameters nor computational complexity. Direct identity-shortcut addition works only when x and F(x) have matching dimensions. If the dimensions differ, the shortcut can use a linear projection to match the dimensions.

Transformers employ a residual connection around each of two sublayers. Highway networks regulate information flow with learned gates. Fixup enables residual networks without normalization.

2 · Why it exists

Deeper neural networks can be more difficult to train.

Deeper is harderThe original ResNet work reports that adding layers to a suitably deep plain model can raise training error.
Preserve a pathIdentity shortcuts add neither parameters nor computational complexity.
Learn the changeThe transformed branch learns a residual relative to the block input.
3 · How it works

Follow one tensor through the shortcut and learned branch.

Element-wise addition forms F(x) + x.
  1. 1 · splitThe shortcut connection adds x to the residual mapping F(x).
  2. 2 · transformStacked layers compute a residual mapping F(x).
  3. 3 · carryDirect identity-shortcut addition requires x and F(x) to have equal dimensions.
  4. 4 · addElement-wise addition forms y = F(x) + x.

When dimensions differ, the shortcut can apply a linear projection before addition.

4 · Where it's used
WhoWhat they askWhat it works with
Vision engineer“How can this block refine a feature map without discarding it?”An identity shortcut around convolutional layers
Language-model engineer“Where does each Transformer sublayer return its input?”Residual connections around Transformer sublayers
Optimisation researcher“Can a deep residual network train without normalization?”Fixup initialization
5 · What it solves, and what it doesn't
solves
  • Residual learning was introduced to ease the optimization of substantially deeper networks.
  • An identity shortcut adds neither parameters nor computational complexity.
  • Identity mappings allow forward and backward signals to pass directly between residual blocks.
doesn't solve
  • Direct identity-shortcut addition requires x and F(x) to have equal dimensions.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. paperDeep Residual Learning for Image Recognition, He, Zhang, Ren and Sun · read 28 Sept 2026
  2. paperIdentity Mappings in Deep Residual Networks, He et al. · read 28 Sept 2026
  3. paperHighway Networks, Srivastava, Greff and Schmidhuber · read 28 Sept 2026
  4. paperAttention Is All You Need, Vaswani et al. · read 28 Sept 2026
  5. paperFixup Initialization: Residual Learning Without Normalization, Zhang, Dauphin and Ma · read 28 Sept 2026