Concepts

Gradient descent

3 min readintermediateUpdated 28 Sept 2026
1 · In one line

Gradient descent repeatedly measures how loss changes with each parameter and moves the parameters a small step toward lower loss.

1 · What it is

Gradient descent is an optimizer, an algorithm that changes model parameters to reduce an objective. The gradient tells how the current loss changes when each parameter changes. Picture the loss as a curve over one weight. The slope at the current weight says which way is downhill.

Each update subtracts the learning rate times the gradient from each weight. A negative slope means the loss falls as the weight grows, so subtracting it raises the weight; a positive slope lowers it. The loop then recomputes the loss at the new weights and repeats. Backpropagation and automatic differentiation produce the gradients; gradient descent uses them to make the update. Libraries such as PyTorch ship it as a ready-made optimiser, with momentum as an option.

The learning rate sets how long each step is. Pick it too small and you wait a long time for the loss to bottom out. Pick it too big and the weight keeps jumping over the low point, landing first on one side and then the other. For a linear model the loss curve is a simple bowl with one lowest point, so the downhill walk does get there. A deep network’s loss surface never has that simple bowl shape. Even so, the walk usually ends somewhere good in practice, though nothing promises it is the very lowest point. Momentum carries part of the previous step into the next one, a bit like a ball that keeps rolling. Adam gives every weight its own step size, worked out from statistics of the gradients that weight has seen so far.

2 · Why it exists

A model may have many parameters, so training needs a repeatable way to search for values with lower loss.

No map of the terrainThe model can only measure the loss at the weights it has right now, so it has to feel its way down.
Local directionA gradient gives the slope of loss with respect to each trainable variable at the current point.
Step controlThe learning rate scales how far each update moves along the negative gradient.
3 · How it works

Follow one weight through a single update.

Illustrative numbers. Each step moves the weight against its slope, scaled by the learning rate, so steps shrink near the bottom.
  1. 1 · measureUse the current weights to make predictions and work out the loss, over the whole dataset or, in stochastic versions, over a small batch.
  2. 2 · slopeBackpropagation finds the gradient, which says how the loss changes as each weight changes.
  3. 3 · stepSubtract the learning rate times the gradient, so each weight moves against its slope.
  4. 4 · repeatMeasure again at the new weights, on the next batch if you use batches, and keep going until the loss stops falling or another stopping rule fires.

The gradient chooses the direction; the learning rate chooses the step size.

4 · Where it's used
WhoWhat they askWhat it works with
Neural-network trainer“How should every weight change after this batch?”Gradients and optimizer update
Regression analyst“Which weight and bias lower mean squared error?”Descent path on the loss surface
Training engineer“Why is loss oscillating or diverging?”Learning rate and loss curve
5 · What it solves, and what it doesn't
solves
  • Gradient descent supplies an iterative rule for finding parameters with lower loss.
  • Automatic differentiation can compute gradients by traversing recorded operations backward.
  • Stochastic versions update after a small subset of examples, so each step need not wait for the whole dataset.
doesn't solve
  • Gradient descent does not choose the loss function or decide whether the training objective matches the product goal.
  • A poor learning rate can make progress crawl, or make the weights overshoot back and forth without settling.
  • Deep networks do not have a single bowl-shaped loss, so a downhill path is not guaranteed to reach the lowest possible loss.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. docsLinear regression: Gradient descent, Google for Developers · read 27 Sept 2026
  2. docsSGD, Keras · read 27 Sept 2026
  3. docsIntroduction to gradients and automatic differentiation, TensorFlow · read 27 Sept 2026
  4. docsSGD, PyTorch · read 27 Sept 2026
  5. paperAdam: A Method for Stochastic Optimization, Kingma and Ba · read 27 Sept 2026
  6. docsLinear regression: Hyperparameters, Google for Developers · read 27 Sept 2026
  7. docsMachine Learning Glossary, Google for Developers · read 27 Sept 2026