Concepts

Learning rate

3 min readbeginnerUpdated 28 Sept 2026
1 · In one line

The learning rate multiplies a gradient, and that product is how far the weight moves.

1 · What it is

Gradient descent knows which way to move each weight, but the learning rate supplies how far. In plain stochastic gradient descent, the new weight is the old weight minus the learning rate times the gradient. With a gradient of 2.5 and a rate of 0.01, the weight moves by 0.025. Double the rate and the same gradient moves the weight twice as far.

Libraries start from different defaults. Keras sets its SGD optimiser to 0.01 by default. PyTorch’s SGD uses 1e-3, which is 0.001. Neither is right everywhere, because the ideal rate depends on the problem. Of all the settings you choose before training, this is often the one that matters most, so it deserves careful tuning.

A bad choice shows up in the loss curve. If the rate is too large, the average loss goes up instead of down. Somewhat too high, and the weights keep overshooting the best values without ever settling. Too small, and each step barely moves, so training takes ages to settle. A useful rule of thumb: the best value usually lies within a factor of two of the highest rate at which training still does not blow up. Libraries also let you shrink the rate as training goes on. In scikit-learn’s SGD models, the same number sets how far each update travels, and PyTorch’s StepLR lowers the rate by a fixed factor every set number of epochs.

2 · Why it exists

A gradient gives a direction, and the rate still has to set the distance.

Loss can riseA rate that is too large can make the average loss increase.
Progress is slowToo small, and every step barely moves the weights, so training crawls.
No shared bestThe right value shifts from one model and dataset to the next.
3 · How it works

Take one gradient through two different rates.

The highlighted step multiplies one gradient by the rate and subtracts. A larger rate makes a longer step.
  1. 1 · readStart from the current weight and the gradient of the loss for that weight.
  2. 2 · scaleMultiply the gradient by the learning rate.
  3. 3 · subtractSubtract that product from the weight to get the next weight.
  4. 4 · compareKeep the gradient fixed and raise the rate, and the step gets longer.

The rate changes the length of the step. The next weight is still the old weight minus the rate times the gradient.

4 · Where it's used
WhoWhat they askWhat it works with
Training engineer“How far should this weight move for the current gradient?”The learning rate set on the optimizer
Run debugger“Why did the loss jump on the last update?”Whether the rate was large enough to raise the loss
Model tuner“Will a smaller rate settle where this one jumps?”The same training loss with a lower rate
Schedule author“Should this rate stay fixed for the whole run?”A later epoch that uses a smaller rate
5 · What it solves, and what it doesn't
solves
  • It turns a gradient, which gives only a direction and steepness, into an actual change in each weight.
  • One number lets you trade speed against stability for the whole run.
  • It can be scheduled to shrink as training goes on.
doesn't solve
  • It does not choose the direction. The sign of the gradient does that.
  • It does not reveal its own best value. You find that by trying rates on your model and data.
  • With momentum switched on, it is no longer the whole step. Past gradients add to each move.
  • It does not change itself during training; a schedule has to do that.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. docsSGD, PyTorch · read 27 Sept 2026
  2. docstf.keras.optimizers.SGD, TensorFlow · read 27 Sept 2026
  3. paperPractical Recommendations for Gradient-Based Training of Deep Architectures, Yoshua Bengio · read 27 Sept 2026
  4. docsStepLR, PyTorch · read 27 Sept 2026
  5. docs1.5. Stochastic Gradient Descent, scikit-learn · read 27 Sept 2026
  6. docsLinear regression: Hyperparameters, Google for Developers · read 27 Sept 2026