Concepts

Layer normalization

3 min readintermediateUpdated 28 Sept 2026
1 · In one line

Layer normalization normalizes selected activations independently for each example in a batch.

1 · What it is

Layer normalization operates on selected features of one example at a time. It computes their mean and variance, subtracts the mean, and divides by the square root of the variance plus epsilon.

The normalized vector then passes through a learned affine transform. Gamma and beta are learnable parameters applied after normalization.

Unlike batch normalization, LayerNorm does not depend on other examples in the mini-batch. It uses statistics from its input in both training and evaluation. The original Transformer placed layer normalization after each residual connection, and later work found that the location of layer normalization matters.

2 · Why it exists

Layer normalization uses statistics from each example instead of across a batch.

Per-example statisticsIt computes the mean and variance of the selected activations for each example independently.
Batch dependenceBatch normalization uses statistics across a mini-batch, so its result depends on which examples share that batch.
Sequence useIn recurrent networks, normalization statistics can be computed separately at each time step.
3 · How it works

Normalize one four-feature vector.

Layer normalization uses statistics from the selected features of each example, not from other examples in the batch.
  1. 1 · selectThe layer selects the configured feature dimensions within one example.
  2. 2 · measureIt computes their mean and variance.
  3. 3 · normalizeIt subtracts the mean and divides by the square root of variance plus a small epsilon.
  4. 4 · adjustIt applies learned scale and bias parameters to the normalized values.

The same input-dependent statistics are used in both training and evaluation modes.

4 · Where it's used
WhoWhat they askWhat it works with
Transformer engineer“Where should normalization sit around this residual block?”Activations before or after attention and feed-forward sublayers
Recurrent model trainer“Can each time step be normalized without batch statistics?”Hidden activations for one sequence at one time step
Framework user“Which dimensions will this LayerNorm call normalize?”The configured normalized shape or axis list
Model debugger“Why do training and evaluation return the same normalization rule?”Whether statistics come from the current input or stored running averages
5 · What it solves, and what it doesn't
solves
  • It normalizes selected features independently for each example in a batch.
  • It uses current-input statistics in both training and evaluation.
  • Learned scale and bias let the layer restore or reshape useful feature ranges.
doesn't solve
  • It does not make every feature independent; mean and variance summarize the group being normalized.
  • It does not use running population statistics like batch normalization.
  • The configured axis or axes determine which dimensions are normalized.
  • Its location in a Transformer residual block matters.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. paperLayer Normalization, Ba, Kiros and Hinton · read 27 Sept 2026
  2. docsLayerNorm, PyTorch · read 27 Sept 2026
  3. docstf.keras.layers.LayerNormalization, TensorFlow · read 27 Sept 2026
  4. paperAttention Is All You Need, Vaswani et al. · read 27 Sept 2026
  5. paperOn Layer Normalization in the Transformer Architecture, Xiong et al. · read 27 Sept 2026