Stochastic gradient descent
Stochastic gradient descent steps the weights using a slope estimated from one randomly picked example.
Gradient descent trains a model by stepping its weights downhill on the training error. Getting that slope exactly means scoring every training example first. Stochastic gradient descent is a drastic shortcut. It guesses the slope from a single randomly picked example and steps straight away. A learning rate sets how big each step is.
The shortcut has a price. Each guess is rough, and that roughness puts a ceiling on how quickly training can settle. Near the bottom, the noise keeps the weights from settling exactly on the minimum. One remedy grows the batch over time, with few examples per step early and more later. Keras can instead pool the gradients from several tiny batches into one update, which calms the noise. Scikit-learn warns to shuffle the training data before fitting.
PyTorch ships it as an optimizer, with momentum as an option. TensorFlow lists the same optimizer as gradient descent with momentum. With momentum set to zero, it is plain gradient descent.
The exact slope needs every example scored before a single step.
Follow one drawn example through the weight update.
- 1 · drawOne example is picked at random from the training set for this iteration.
- 2 · scoreThe error on that one example gives an estimate of the true gradient.
- 3 · stepThe next weight equals the current weight minus a gain times that example's gradient.
- 4 · repeatThe loop starts again with a fresh draw, and new data can join as it comes in.
The weight step uses only this draw, and the rest of the training set waits.
| Who | What they ask | What it works with |
|---|---|---|
| Vision lab | “Can we update the network without scoring every photo?” | One random photo from the training folders |
| Spam filter | “How do the weights change when a new message arrives?” | The loss on that single message |
| On-device model | “Can training continue as fresh taps come in?” | Examples handled as they arrive |
| Tuning lead | “Why did the loss jump on the last step?” | The one example that was drawn |
- Each step estimates the training gradient from one randomly picked example.
- Training can use data the moment it arrives, with no record kept of past picks.
- It usually reaches a good model sooner than batch learning, and the gap is largest on big data sets full of repeats.
- It is a way to train many kinds of model, not a model of its own.
- Each step follows a rough guess, so its direction can wander off the true slope.
- The same noise stops the weights from settling exactly at the minimum.
- Set the learning rate too high and each step can land farther from the minimum instead of closer.
- It needs several hand-tuned settings, and inputs on very different scales can throw it off.
Sources used
This explainer is written in original language. The links below support its factual claims.
- paperLarge-Scale Machine Learning with Stochastic Gradient Descent, Bottou, COMPSTAT 2010 · read 27 Sept 2026
- paperEfficient BackProp, LeCun, Bottou, Orr, and Müller, 1998 · read 27 Sept 2026
- docs1.5. Stochastic Gradient Descent, scikit-learn · read 27 Sept 2026
- docsSGD, PyTorch · read 27 Sept 2026
- docsSGD, Keras · read 27 Sept 2026
- docstf.keras.optimizers.SGD, TensorFlow · read 27 Sept 2026