Concepts

Residual networks

3 min readintermediateUpdated 28 Sept 2026
1 · In one line

A residual block adds a learned residual function F(x) to a shortcut carrying the block input x.

1 · What it is

The original mapping is recast as F(x) + x. The stacked layers learn the residual mapping F(x). A shortcut connection and element-wise addition perform the F(x) + x operation.

Residual learning addressed the degradation problem observed in deeper plain networks. Identity shortcut connections add neither extra parameters nor computational complexity. When dimensions differ, the shortcut can use a linear projection.

Wide residual networks multiply the number of convolutional features by a factor k. ResNeXt calls the size of the set of transformations cardinality. Fixup enables very deep residual networks to train without normalization.

2 · Why it exists

Adding layers to a suitably deep plain network can increase its training error.

DegradationThe original ResNet paper calls this increase in training error the degradation problem.
Direct pathIdentity skip connections let forward and backward signals travel directly between residual blocks.
Additive targetThe central idea of ResNets is to learn an additive residual function F.
3 · How it works

Follow one activation through a residual block.

A residual block with an identity shortcut Input x splits into a learned residual branch and an identity shortcut. Two learned layers produce F of x. An element-wise addition combines F of x with x to form output y. A note explains projection shortcuts for changed dimensions.
The main branch learns F(x). The shortcut carries x, and element-wise addition forms F(x) + x.
  1. 1 · shortcutA shortcut connection skips one or more layers.
  2. 2 · transformThe stacked layers in the residual branch compute F(x).
  3. 3 · carryIdentity shortcut connections add neither extra parameters nor computational complexity.
  4. 4 · addElement-wise addition combines F(x) and x.

If the dimensions differ, the shortcut can use a learned linear projection before addition.

4 · Where it's used
WhoWhat they askWhat it works with
Vision researcher“Can a deeper image classifier avoid the plain network's degradation problem?”Residual blocks in a convolutional network
Architecture researcher“What happens if each residual stage is widened?”Feature width multiplier
Optimisation researcher“Can a residual network train without normalization?”Fixup initialization rules
5 · What it solves, and what it doesn't
solves
  • Residual learning addressed the degradation problem observed in deeper plain networks.
  • Identity shortcuts add no extra parameters or computational complexity.
  • Identity mappings provide a direct route for forward and backward signals between blocks.
doesn't solve
  • The tensors F(x) and x must have equal dimensions for direct element-wise addition.
  • The original formulation uses a projection shortcut when dimensions change.
  • In Wide ResNets, parameter count and computational complexity are quadratic in the width multiplier.
  • Fixup was introduced because standard initialization without normalization can cause exploding gradients in residual networks.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. paperDeep Residual Learning for Image Recognition, He, Zhang, Ren and Sun · read 28 Sept 2026
  2. paperIdentity Mappings in Deep Residual Networks, He et al. · read 28 Sept 2026
  3. paperWide Residual Networks, Zagoruyko and Komodakis · read 28 Sept 2026
  4. paperAggregated Residual Transformations for Deep Neural Networks, Xie et al. · read 28 Sept 2026
  5. paperFixup Initialization, Zhang, Dauphin and Ma · read 28 Sept 2026