Concepts

Epoch

4 min readbeginnerUpdated 28 Sept 2026
1 · In one line

An epoch is one complete pass through every example used to train a model.

1 · What it is

An epoch is one full pass through the training set, so each example is processed once. In Keras fit, that pass is an iteration over the inputs and targets, unless steps_per_epoch ends it sooner. The Keras FAQ calls the same cutoff one pass over the entire dataset. A training batch produces only one update to the model. PyTorch’s beginner tutorial counts epochs as the number of times to iterate over the dataset. Batch size there is how many samples pass through before the parameters, the numbers inside the model, are updated. A step is that one update.

Google’s glossary writes one epoch as N divided by the batch size. One thousand examples and a batch of 50 make 20 iterations. With 1,000 examples and batches of 100, one epoch takes 10 iterations. The same 1,000 examples and 20 epochs update the weights 20 times when the batch is the full set. Those 20 epochs update the weights 200 times when the batch size is 100. When steps_per_epoch is left unset, Keras divides the number of samples by the batch size. The figure applies that division to eight examples in batches of two, which is four updates. For scikit-learn’s sgd and adam solvers, max_iter counts epochs, not gradient steps.

Training typically requires many epochs. Google’s course says that, in general, more epochs produce a better model, but they also take more time. That is not a guarantee. For NumPy arrays, with shuffle left at its default, Keras randomly reorders the training data before each epoch. Validation accuracy can peak after a number of epochs and then stagnate or start decreasing. Training too long makes the model learn patterns from the training data that do not generalize to test data. The model may then fit the training data so closely that it predicts new examples poorly.

2 · Why it exists

Twenty epochs over 1,000 examples become 200 weight updates when the batch holds 100.

A step is smallerOne batch makes a single update. An epoch is the pass over the whole set, unless steps_per_epoch ends it after a fixed number of batches.
Size changes the countThe same number of epochs can hide a very different number of updates when the batch size changes.
Order is not fixedFor NumPy arrays, Keras randomly shuffles the training rows before each epoch when shuffle is left on.
3 · How it works

Follow eight examples until the epoch counter moves.

Eight examples in batches of two make four updates. The dark box is the last update, when the epoch counter advances.
  1. 1 · groupSplit the eight examples into four batches of two.
  2. 2 · updateEach batch is one step, and that step makes one update to the model.
  3. 3 · sweepFour updates use every example in the set once.
  4. 4 · advanceWhen the fourth batch finishes, the epoch counter moves from zero to one.

The epoch is the full sweep. The counter moves only after the last batch.

4 · Where it's used
WhoWhat they askWhat it works with
Tutorial reader“How many times does this loop see the whole dataset?”The epochs value in the training script
Keras developer“Why did this epoch stop before every row was used?”The steps_per_epoch argument on fit
scikit-learn user“Is max_iter a gradient step or a full pass?”The note for the sgd and adam solvers
Training lead“Did a smaller batch add updates or add passes?”The epoch count next to the step count
5 · What it solves, and what it doesn't
solves
  • It marks one complete pass, and Keras runs evaluation at the end of every epoch.
  • It separates that pass from a step, which is one batch update.
  • It states how many times each training example is used.
  • It shows why a smaller batch puts more updates inside the same epoch.
doesn't solve
  • Finishing the epoch you set is not proof the model has learned.
  • The same epoch count is not the same number of weight updates when the batch size changes.
  • Shuffling can change the order in which that pass meets the examples.
  • Training for more epochs can fit the training data so closely that results on new examples get worse.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. docstf.keras.Model, TensorFlow · read 27 Sept 2026
  2. docsKeras FAQ, Keras · read 27 Sept 2026
  3. docsOptimizing Model Parameters, PyTorch · read 27 Sept 2026
  4. docsMLPClassifier, scikit-learn · read 27 Sept 2026
  5. docsLinear regression: Hyperparameters, Google for Developers · read 27 Sept 2026
  6. docsMachine Learning Glossary, Google for Developers · read 27 Sept 2026
  7. docsOverfit and underfit, TensorFlow · read 27 Sept 2026