Concepts

Backpropagation

4 min readintermediateUpdated 28 Sept 2026
1 · In one line

Backpropagation computes each parameter's gradient by applying the chain rule from the last layer back to the first.

1 · What it is

A neural network makes a prediction by pushing numbers forward through its layers. A loss scores how far that prediction landed from the right answer. Backpropagation then asks a question about every weight: if this one moved a little, how much would the loss change? That number is the weight’s gradient.

The trick is to start at the end. The calculation starts at the last layer and works back to the first. The chain rule from calculus lets each layer combine the slope handed back to it with its own local slope, then pass the product one layer further down. You rarely write this yourself. TensorFlow records the forward operations and walks them in reverse to get the gradients. PyTorch does the same for any function that ends in a single number.

The idea is older than its fame. Backpropagation is a special case of a wider family of techniques called automatic differentiation. Surveys often credit Linnainmaa’s 1970 work as the earliest account of the reverse method in print. In 1974 Werbos wrote a formal version for systems that move in discrete time steps. Machine-learning researchers reinvented it several times. A 1986 Nature paper by Rumelhart, Hinton and Williams made it famous for training neural networks. That paper showed hidden units learning useful features of the task on their own.

On this page, backpropagation only measures. In a library such as PyTorch, the backward pass fills in the gradients, and only afterwards does an optimiser step use them to move the weights. Backpropagation is what makes gradient descent practical for networks with many layers. Some guides use the word backpropagation for the whole training procedure, weight updates included. Backpropagation also has a weak spot. In deep networks, gradients for the layers near the input can shrink towards zero. Those layers then barely learn, or stop learning. With very large weights, the gradients can instead grow too large for training to settle.

2 · Why it exists

A network's error shows up at the output, but the weights that caused it sit in every layer.

One score, many weightsTraining ends each example with a single loss number, yet every weight in every layer had a hand in it.
Hidden layers get no answerOnly the output can be compared with a target. The layers in between have no right answer of their own.
Descent needs a slopeGradient descent can move a weight only once it knows which direction lowers the loss.
3 · How it works

Follow one example forward, then carry the blame back layer by layer.

The backward row sits under the layer it serves and runs right to left. The chain rule repeats at every layer; the dark box shows it at the hidden layer.
  1. 1 · forwardThe network runs one example forward, and the library records each operation in order.
  2. 2 · scoreA loss compares the output with the target and gives one number.
  3. 3 · startThe slope of the loss is found first, at the last layer.
  4. 4 · multiplyEach layer multiplies the slope it receives by its own local slope, which is the chain rule, and passes the result back.
  5. 5 · collectWhen the sweep reaches the first layer, every weight holds a gradient and none has changed yet.

Here backpropagation means the step that returns gradients. In code, a separate optimiser step is what changes the weights.

4 · Where it's used
WhoWhat they askWhat it works with
Image model“Which way should this filter weight move to cut the error?”The loss on one labelled photo
Speech model“How did this layer change the error on this recording?”The loss on the recording and its transcript
Language model“Which weights raised the loss on this sentence?”The loss against the next-word target
Library author“Can this new layer send a gradient backward?”Whether each operation in the layer has a derivative
5 · What it solves, and what it doesn't
solves
  • One backward sweep gives every weight its own gradient.
  • It reaches hidden layers, which have no target of their own.
  • It works on any chain of differentiable steps, so libraries can run it for you.
doesn't solve
  • It changes no weights. An optimiser does that, using the gradients.
  • It cannot choose the learning rate, the number of layers or the training data.
  • In deep networks, early layers can receive gradients so small that they barely learn.
  • An operation with no gradient defined breaks the backward sweep.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. paperLearning representations by back-propagating errors, Rumelhart, Hinton, and Williams, Nature · read 27 Sept 2026
  2. docsAutomatic differentiation package - torch.autograd, PyTorch · read 27 Sept 2026
  3. docsIntroduction to gradients and automatic differentiation, TensorFlow · read 27 Sept 2026
  4. paperEfficient BackProp, LeCun, Bottou, Orr, and Müller · read 27 Sept 2026
  5. docsNeural networks: Training using backpropagation, Google for Developers · read 27 Sept 2026
  6. paperAutomatic differentiation in machine learning: a survey, Baydin, Pearlmutter, Radul, and Siskind (arXiv; JMLR) · read 27 Sept 2026
  7. docstorch.optim, PyTorch · read 27 Sept 2026