Concepts

Batch size

3 min readbeginnerUpdated 28 Sept 2026
1 · In one line

Batch size is how many training examples are grouped before the weights change once.

1 · What it is

Batch size is how many training examples are grouped before the weights update once. In Keras the default is 32 examples per update. PyTorch’s data loader starts at 1.

A size of one uses a single random example. SGD skips the exact gradient over the whole training set, so each step follows a rough, noisy guess of it. A gradient says how the loss changes if a weight moves. In mini-batch training, the group holds more than one example but fewer than the whole dataset. Each example gives its own gradient; these are averaged, and only then do the weights and bias change, once for the whole group. Waiting for the whole dataset before each change gets impractical, because real datasets can reach millions of examples. The smaller the group, the closer training gets to one-example SGD; the larger it is, the closer it gets to full-batch descent.

A bigger group needs more graphics-processor memory to hold its inputs, activations (values inside the model) and gradients. The learning rate is a separate setting that controls how far the weights move. Pick a different batch size and the learning rate that worked before may need retuning. Take 1,000 training examples split into groups of 100. That makes 10 iterations, and so 10 weight updates, before the model has processed each example once, which is one epoch.

2 · Why it exists

A minibatch sits between one example and the full training set.

One example wobblesA size of one uses a single random example. That step is very noisy. The loss can rise on the step.
All or oneFull-batch gradient descent reads every example before each step. Plain SGD reads one at a time.
Memory fills upMore examples per batch means more graphics-processor memory for their inputs, activations and gradients. Too big a batch is a common cause of out-of-memory crashes.
3 · How it works

Follow four examples from a set of eight into one weight step.

The batch's gradients are averaged into one weight step. A size of one would step after each example.
  1. 1 · groupThe data loader pulls out as many examples as the batch size says and passes them on as one group.
  2. 2 · scoreEach example yields a gradient, how the loss changes if a weight moves.
  3. 3 · averageThe group's gradients are combined into one average, and that average sets a single change to the weights.
  4. 4 · moveThe learning rate chooses how far that one update moves.

One batch, one update. A size of 1 steps after every example instead.

4 · Where it's used
WhoWhat they askWhat it works with
Student on a first run“Why did the loss rise on this step?”Whether the batch size was one
Engineer with a small graphics processor“Why did training run out of memory?”How many examples the batch tried to store
Scientist comparing two runs“Can a very large batch miss a flatter minimum?”Results on large-batch training
Engineer setting the loader“What happens to the leftover examples?”Whether the last incomplete batch is kept
5 · What it solves, and what it doesn't
solves
  • It names how many examples are used before the weights update once.
  • A larger group gives a steadier gradient estimate than a smaller one.
  • The group can be kept small enough to fit in graphics-processor memory.
  • One group still produces a single weight update, even when it holds many examples.
doesn't solve
  • Batch size is one setting among several, next to learning rate and epoch count. Pick a new batch size and the old learning rate may stop working, so retune it.
  • The last batch of a pass can be smaller when the example count does not divide evenly, unless the loader drops it.
  • scikit-learn's L-BFGS solver does not train with minibatches.
  • In one study, very large batches kept landing in sharp dips of the loss, and sharp dips predict worse results on new data. Small batches found flatter dips.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. docstorch.utils.data — PyTorch 2.14 documentation, PyTorch · read 27 Sept 2026
  2. docsModel training APIs, Keras · read 27 Sept 2026
  3. paperStochastic Gradient Descent Tricks, Léon Bottou · read 27 Sept 2026
  4. docsMLPClassifier — scikit-learn 1.9.1 documentation, scikit-learn · read 27 Sept 2026
  5. docs1.5. Stochastic Gradient Descent — scikit-learn 1.9.1 documentation, scikit-learn · read 27 Sept 2026
  6. docsLinear regression: Hyperparameters | Machine Learning | Google for Developers, Google for Developers · read 27 Sept 2026
  7. docsData Loading Optimization in PyTorch — PyTorch Tutorials 2.14.0+cu130 documentation, PyTorch · read 27 Sept 2026
  8. docs1.17. Neural network models (supervised) — scikit-learn 1.9.1 documentation, scikit-learn · read 27 Sept 2026
  9. paperOn Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minima, Keskar et al., ICLR 2017 · read 27 Sept 2026