Rectified Linear Unit
ReLU is an activation function that replaces every negative input with zero and leaves every positive input unchanged.
ReLU acts on one number at a time. For an input below zero, its output is zero. For an input above zero, its output equals the input. This piecewise rule is written as max(0, x).
That bend at zero matters. Without a non-linear activation, stacking linear layers still yields a linear function. ReLU lets a deep network build non-linear mappings while remaining cheap to evaluate. Rectifier networks can also produce many exact zeros, so only part of a layer may be active for a given example.
ReLU is not differentiable at exactly zero. Rectifier-aware initialization was developed for training very deep rectified models from scratch.
Neural networks need non-linear steps between their linear layers.
Follow five pre-activation values through ReLU.
- 1 · receiveAn activation function transforms the output of each node in a layer.
- 2 · compareReLU compares that value with zero.
- 3 · rectifyIt returns zero for a negative input and returns the input itself otherwise.
- 4 · passThe result becomes an input to the next layer.
ReLU is written as max(0, x).
| Who | What they ask | What it works with |
|---|---|---|
| Vision engineer | “Which hidden features should stay active for this image?” | Convolution outputs before the next layer |
| Tabular model team | “How can stacked dense layers learn non-linear boundaries?” | Weighted sums inside a feed-forward network |
| Model debugger | “Why is this unit always returning zero?” | The unit's pre-activation values across training examples |
| Architecture researcher | “Which activation should follow a linear layer?” | Gradient flow, sparsity and task results |
- It inserts a non-linear transform between linear layers.
- It is simple to compute because it only compares the input with zero.
- Negative inputs produce exact zeros, which creates sparse activations.
- ReLU is not differentiable at exactly zero.
- It does not bound positive activations; positive inputs pass through unchanged.
Sources used
This explainer is written in original language. The links below support its factual claims.
- docsNeural networks: Activation functions, Google for Developers · read 27 Sept 2026
- docsReLU, PyTorch · read 27 Sept 2026
- paperDeep Sparse Rectifier Neural Networks, Glorot, Bordes and Bengio · read 27 Sept 2026
- paperDelving Deep into Rectifiers, He, Zhang, Ren and Sun · read 27 Sept 2026
- docsBuild the Neural Network, PyTorch · read 27 Sept 2026