Learning rate
The learning rate multiplies a gradient, and that product is how far the weight moves.
Gradient descent knows which way to move each weight, but the learning rate supplies how far. In plain stochastic gradient descent, the new weight is the old weight minus the learning rate times the gradient. With a gradient of 2.5 and a rate of 0.01, the weight moves by 0.025. Double the rate and the same gradient moves the weight twice as far.
Libraries start from different defaults. Keras sets its SGD optimiser to 0.01 by default. PyTorch’s SGD uses 1e-3, which is 0.001. Neither is right everywhere, because the ideal rate depends on the problem. Of all the settings you choose before training, this is often the one that matters most, so it deserves careful tuning.
A bad choice shows up in the loss curve. If the rate is too large, the average loss goes up instead of down. Somewhat too high, and the weights keep overshooting the best values without ever settling. Too small, and each step barely moves, so training takes ages to settle. A useful rule of thumb: the best value usually lies within a factor of two of the highest rate at which training still does not blow up. Libraries also let you shrink the rate as training goes on. In scikit-learn’s SGD models, the same number sets how far each update travels, and PyTorch’s StepLR lowers the rate by a fixed factor every set number of epochs.
A gradient gives a direction, and the rate still has to set the distance.
Take one gradient through two different rates.
- 1 · readStart from the current weight and the gradient of the loss for that weight.
- 2 · scaleMultiply the gradient by the learning rate.
- 3 · subtractSubtract that product from the weight to get the next weight.
- 4 · compareKeep the gradient fixed and raise the rate, and the step gets longer.
The rate changes the length of the step. The next weight is still the old weight minus the rate times the gradient.
| Who | What they ask | What it works with |
|---|---|---|
| Training engineer | “How far should this weight move for the current gradient?” | The learning rate set on the optimizer |
| Run debugger | “Why did the loss jump on the last update?” | Whether the rate was large enough to raise the loss |
| Model tuner | “Will a smaller rate settle where this one jumps?” | The same training loss with a lower rate |
| Schedule author | “Should this rate stay fixed for the whole run?” | A later epoch that uses a smaller rate |
- It turns a gradient, which gives only a direction and steepness, into an actual change in each weight.
- One number lets you trade speed against stability for the whole run.
- It can be scheduled to shrink as training goes on.
- It does not choose the direction. The sign of the gradient does that.
- It does not reveal its own best value. You find that by trying rates on your model and data.
- With momentum switched on, it is no longer the whole step. Past gradients add to each move.
- It does not change itself during training; a schedule has to do that.
Sources used
This explainer is written in original language. The links below support its factual claims.
- docsSGD, PyTorch · read 27 Sept 2026
- docstf.keras.optimizers.SGD, TensorFlow · read 27 Sept 2026
- paperPractical Recommendations for Gradient-Based Training of Deep Architectures, Yoshua Bengio · read 27 Sept 2026
- docsStepLR, PyTorch · read 27 Sept 2026
- docs1.5. Stochastic Gradient Descent, scikit-learn · read 27 Sept 2026
- docsLinear regression: Hyperparameters, Google for Developers · read 27 Sept 2026