Layer normalization
Layer normalization normalizes selected activations independently for each example in a batch.
Layer normalization operates on selected features of one example at a time. It computes their mean and variance, subtracts the mean, and divides by the square root of the variance plus epsilon.
The normalized vector then passes through a learned affine transform. Gamma and beta are learnable parameters applied after normalization.
Unlike batch normalization, LayerNorm does not depend on other examples in the mini-batch. It uses statistics from its input in both training and evaluation. The original Transformer placed layer normalization after each residual connection, and later work found that the location of layer normalization matters.
Layer normalization uses statistics from each example instead of across a batch.
Normalize one four-feature vector.
- 1 · selectThe layer selects the configured feature dimensions within one example.
- 2 · measureIt computes their mean and variance.
- 3 · normalizeIt subtracts the mean and divides by the square root of variance plus a small epsilon.
- 4 · adjustIt applies learned scale and bias parameters to the normalized values.
The same input-dependent statistics are used in both training and evaluation modes.
| Who | What they ask | What it works with |
|---|---|---|
| Transformer engineer | “Where should normalization sit around this residual block?” | Activations before or after attention and feed-forward sublayers |
| Recurrent model trainer | “Can each time step be normalized without batch statistics?” | Hidden activations for one sequence at one time step |
| Framework user | “Which dimensions will this LayerNorm call normalize?” | The configured normalized shape or axis list |
| Model debugger | “Why do training and evaluation return the same normalization rule?” | Whether statistics come from the current input or stored running averages |
- It normalizes selected features independently for each example in a batch.
- It uses current-input statistics in both training and evaluation.
- Learned scale and bias let the layer restore or reshape useful feature ranges.
- It does not make every feature independent; mean and variance summarize the group being normalized.
- It does not use running population statistics like batch normalization.
- The configured axis or axes determine which dimensions are normalized.
- Its location in a Transformer residual block matters.
Sources used
This explainer is written in original language. The links below support its factual claims.
- paperLayer Normalization, Ba, Kiros and Hinton · read 27 Sept 2026
- docsLayerNorm, PyTorch · read 27 Sept 2026
- docstf.keras.layers.LayerNormalization, TensorFlow · read 27 Sept 2026
- paperAttention Is All You Need, Vaswani et al. · read 27 Sept 2026
- paperOn Layer Normalization in the Transformer Architecture, Xiong et al. · read 27 Sept 2026