Concepts

Stochastic gradient descent

3 min readbeginnerUpdated 28 Sept 2026
1 · In one line

Stochastic gradient descent steps the weights using a slope estimated from one randomly picked example.

1 · What it is

Gradient descent trains a model by stepping its weights downhill on the training error. Getting that slope exactly means scoring every training example first. Stochastic gradient descent is a drastic shortcut. It guesses the slope from a single randomly picked example and steps straight away. A learning rate sets how big each step is.

The shortcut has a price. Each guess is rough, and that roughness puts a ceiling on how quickly training can settle. Near the bottom, the noise keeps the weights from settling exactly on the minimum. One remedy grows the batch over time, with few examples per step early and more later. Keras can instead pool the gradients from several tiny batches into one update, which calms the noise. Scikit-learn warns to shuffle the training data before fitting.

PyTorch ships it as an optimizer, with momentum as an option. TensorFlow lists the same optimizer as gradient descent with momentum. With momentum set to zero, it is plain gradient descent.

2 · Why it exists

The exact slope needs every example scored before a single step.

Slow full passesBatch learning scores the whole training set to get the exact slope, and only then moves the weights. If the data holds many repeats, most of that work redoes sums already done.
Noisy single slopesOne example gives only a rough guess of the true slope. So a step can point a little off the steepest way down.
Steps that overshootToo big a learning rate throws a weight past the bottom to a worse spot, and training blows up.
3 · How it works

Follow one drawn example through the weight update.

One randomly picked example supplies the loss. The weight step uses that estimate, and the other examples sit out.
  1. 1 · drawOne example is picked at random from the training set for this iteration.
  2. 2 · scoreThe error on that one example gives an estimate of the true gradient.
  3. 3 · stepThe next weight equals the current weight minus a gain times that example's gradient.
  4. 4 · repeatThe loop starts again with a fresh draw, and new data can join as it comes in.

The weight step uses only this draw, and the rest of the training set waits.

4 · Where it's used
WhoWhat they askWhat it works with
Vision lab“Can we update the network without scoring every photo?”One random photo from the training folders
Spam filter“How do the weights change when a new message arrives?”The loss on that single message
On-device model“Can training continue as fresh taps come in?”Examples handled as they arrive
Tuning lead“Why did the loss jump on the last step?”The one example that was drawn
5 · What it solves, and what it doesn't
solves
  • Each step estimates the training gradient from one randomly picked example.
  • Training can use data the moment it arrives, with no record kept of past picks.
  • It usually reaches a good model sooner than batch learning, and the gap is largest on big data sets full of repeats.
  • It is a way to train many kinds of model, not a model of its own.
doesn't solve
  • Each step follows a rough guess, so its direction can wander off the true slope.
  • The same noise stops the weights from settling exactly at the minimum.
  • Set the learning rate too high and each step can land farther from the minimum instead of closer.
  • It needs several hand-tuned settings, and inputs on very different scales can throw it off.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. paperLarge-Scale Machine Learning with Stochastic Gradient Descent, Bottou, COMPSTAT 2010 · read 27 Sept 2026
  2. paperEfficient BackProp, LeCun, Bottou, Orr, and Müller, 1998 · read 27 Sept 2026
  3. docs1.5. Stochastic Gradient Descent, scikit-learn · read 27 Sept 2026
  4. docsSGD, PyTorch · read 27 Sept 2026
  5. docsSGD, Keras · read 27 Sept 2026
  6. docstf.keras.optimizers.SGD, TensorFlow · read 27 Sept 2026