ReLUConcepts

Rectified Linear Unit

3 min readbeginnerUpdated 28 Sept 2026
1 · In one line

ReLU is an activation function that replaces every negative input with zero and leaves every positive input unchanged.

1 · What it is

ReLU acts on one number at a time. For an input below zero, its output is zero. For an input above zero, its output equals the input. This piecewise rule is written as max(0, x).

That bend at zero matters. Without a non-linear activation, stacking linear layers still yields a linear function. ReLU lets a deep network build non-linear mappings while remaining cheap to evaluate. Rectifier networks can also produce many exact zeros, so only part of a layer may be active for a given example.

ReLU is not differentiable at exactly zero. Rectifier-aware initialization was developed for training very deep rectified models from scratch.

2 · Why it exists

Neural networks need non-linear steps between their linear layers.

Layers stay linearStacking only linear operations still produces a linear function.
Smooth units saturateSigmoid and tanh can be more susceptible to vanishing gradients during training.
Dense activity costsA rectifier can create exact zeros, producing sparse internal representations.
3 · How it works

Follow five pre-activation values through ReLU.

ReLU clips the negative half of the number line and passes the positive half unchanged.
  1. 1 · receiveAn activation function transforms the output of each node in a layer.
  2. 2 · compareReLU compares that value with zero.
  3. 3 · rectifyIt returns zero for a negative input and returns the input itself otherwise.
  4. 4 · passThe result becomes an input to the next layer.

ReLU is written as max(0, x).

4 · Where it's used
WhoWhat they askWhat it works with
Vision engineer“Which hidden features should stay active for this image?”Convolution outputs before the next layer
Tabular model team“How can stacked dense layers learn non-linear boundaries?”Weighted sums inside a feed-forward network
Model debugger“Why is this unit always returning zero?”The unit's pre-activation values across training examples
Architecture researcher“Which activation should follow a linear layer?”Gradient flow, sparsity and task results
5 · What it solves, and what it doesn't
solves
  • It inserts a non-linear transform between linear layers.
  • It is simple to compute because it only compares the input with zero.
  • Negative inputs produce exact zeros, which creates sparse activations.
doesn't solve
  • ReLU is not differentiable at exactly zero.
  • It does not bound positive activations; positive inputs pass through unchanged.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. docsNeural networks: Activation functions, Google for Developers · read 27 Sept 2026
  2. docsReLU, PyTorch · read 27 Sept 2026
  3. paperDeep Sparse Rectifier Neural Networks, Glorot, Bordes and Bengio · read 27 Sept 2026
  4. paperDelving Deep into Rectifiers, He, Zhang, Ren and Sun · read 27 Sept 2026
  5. docsBuild the Neural Network, PyTorch · read 27 Sept 2026