Concepts

Activation function

4 min readbeginnerUpdated 28 Sept 2026
1 · In one line

An activation function is the non-linear step a neuron applies to its weighted sum, so a network of layers can learn curved patterns, not only straight lines.

1 · What it is

A neuron does two things. First it multiplies each input by a weight and adds up the products, giving a weighted sum. Then it runs that sum through an activation function and hands the result, not the raw sum, to the next layer. With sigmoid as the activation, a sum of −2.0 leaves the neuron as roughly 0.12.

That function must be non-linear: no mix of adding and multiplying can reproduce it. Plot it and you always see a bend or a kink somewhere. Why does that matter? Strip the function out and every layer is plain multiply-and-add, so a stack of ten layers collapses into one straight-line rule. Extra layers then buy nothing. Put a bend after each layer and the story changes: every layer builds on the last, so deeper layers can capture richer patterns in the original data.

Sigmoid, tanh and ReLU are three common choices. Sigmoid squeezes any number into the range 0 to 1, and tanh into −1 to 1. ReLU, short for rectified linear unit, outputs 0 for a negative or zero input and passes a positive input through unchanged. Google’s crash course recommends starting with ReLU. In a 2011 study, networks built from rectifiers did as well as tanh networks or better, despite the kink at zero where ReLU has no defined slope. The original Transformer put a ReLU between the two linear layers of each feed-forward block. BERT swapped it for GELU, following OpenAI’s GPT. GELU multiplies an input by the probability that a standard normal value is at or below it. So rather than a hard on-off switch at zero, it scales each input smoothly according to how big it is. Another option, swish, also called SiLU, multiplies an input by its own sigmoid. Meta’s LLaMA models replaced ReLU with SwiGLU. Classifiers often end with softmax, since its outputs can be read as probabilities spread across the classes.

2 · Why it exists

Without a bend between layers, a neural network can only learn straight-line patterns.

Sums stay straightA neuron on its own only multiplies and adds. Feed one straight-line step into another and you still get a straight line.
Depth alone failsPiling on more multiply-and-add layers never produces a curve by itself.
Raw scores lack shapeA raw model output is not a probability until a function such as sigmoid squeezes it into the range 0 to 1.
3 · How it works

Follow one neuron's sum through the activation.

Illustrative sums. The same z comes out differently under each function, and none of the three graphs is a single straight line.
  1. 1 · sumThe neuron multiplies each input by its weight, adds the results and adds a bias.
  2. 2 · bendIt passes that weighted sum through the activation function, a non-linear transform.
  3. 3 · passThe transformed value, not the raw sum, goes on as an input to the next layer.
  4. 4 · learnIn training, gradients flow back through the same function, and where a ReLU outputs 0 they stop.

Linear on linear stays linear. The activation is what lets extra layers add something new.

4 · Where it's used
WhoWhat they askWhat it works with
Photo app team“Is this picture a dog, a cat or a horse?”The final layer's scores, turned into class probabilities by softmax
Email provider“How likely is it that this message is spam?”One output score, squeezed between 0 and 1 by sigmoid
Language model engineers“Which activation should go inside each feed-forward block?”The hidden units of every transformer layer, using ReLU, GELU or SwiGLU
Anyone training a deep network“Why have so many units stopped changing?”ReLU units whose weighted sums stay below 0
5 · What it solves, and what it doesn't
solves
  • With a bend after every layer, a deep network can fit tangled links between what goes in and what comes out.
  • Squeezes a value into a fixed range when one is needed, such as 0 to 1 for a probability.
  • ReLU suffers less from vanishing gradients than sigmoid or tanh, and it is cheaper to calculate.
  • ReLU produces many exact zeros, a sparse pattern that suits naturally sparse data.
doesn't solve
  • Nothing forces one pick. In principle almost any maths function could fill the role, so the designer has to choose.
  • Sigmoid and tanh are more exposed to vanishing gradients than ReLU. When the gradient reaching a layer shrinks toward zero, that layer barely learns.
  • A ReLU unit whose sum falls below 0 can get stuck outputting 0, so it stops learning.
  • It does not prevent exploding gradients. Batch normalization or a lower learning rate help there.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. docsNeural networks: Activation functions (Machine Learning Crash Course), Google for Developers · read 27 Sept 2026
  2. docsNeural networks: Training using backpropagation (Machine Learning Crash Course), Google for Developers · read 27 Sept 2026
  3. docsMachine Learning Glossary, Google for Developers · read 27 Sept 2026
  4. docsLayer activation functions, Keras · read 27 Sept 2026
  5. docsTanh (PyTorch 2.14 documentation), PyTorch · read 27 Sept 2026
  6. paperDeep Sparse Rectifier Neural Networks, Glorot, Bordes and Bengio, AISTATS 2011 (PMLR) · read 27 Sept 2026
  7. paperGaussian Error Linear Units (GELUs), Hendrycks and Gimpel (arXiv) · read 27 Sept 2026
  8. paperAttention Is All You Need, Vaswani et al. (arXiv) · read 27 Sept 2026
  9. paperBERT: Pre-training of Deep Bidirectional Transformers for Language Understanding, Devlin et al. (arXiv) · read 27 Sept 2026
  10. paperLLaMA: Open and Efficient Foundation Language Models, Touvron et al., Meta AI (arXiv) · read 27 Sept 2026