Model training
Training is the repeated loop that nudges a model's internal numbers so its predictions get closer to the right answers in its examples.
Training is how a model goes from useless to useful. A model is a large set of numbers, called parameters or weights, and training is the search for good values. The weights start as small random numbers, so the first answers are mostly wrong. The model then reads examples, often many times over, and adjusts its weights a little after each look. Once training ends, the weights are frozen and the model is used to make predictions, which is called inference.
Each adjustment follows the same loop. The model makes predictions for a small batch of examples. A loss function scores how wrong they were as one number. Backpropagation, popularised for neural networks by a 1986 paper, then works backwards through the model to find each weight’s gradient: which way to move it to lower the loss. Finally an optimiser such as gradient descent, or a variant such as Adam, moves every weight a small step. The learning rate sets the step size: with a gradient of 2.5 and a learning rate of 0.01, a weight moves by 0.025. In Google’s worked example, the loss starts at 303.71 and falls to 42.17 after five weight updates.
One pass through every example is an epoch, and training usually takes many. In PyTorch’s beginner tutorial, with batches of 64, test accuracy climbs from 45.9 per cent after the first epoch to 71.1 per cent after the tenth. Training stops when the loss stops falling, or earlier, when loss on held-out validation data starts to rise, because training too long makes a model overfit.
For large language models, training is a major engineering project. Large models are first pretrained on a huge general dataset, then fine-tuned on hundreds or thousands of task examples. Two published examples show the scale. GPT-3, described in a 2020 paper, has 175 billion parameters and was pretrained on 300 billion tokens using V100 graphics processors (GPUs); that run took several thousand petaflop/s-days of compute, against tens for GPT-2. One petaflop/s-day is 8.64 × 10^19 calculations. Meta’s Llama 3.1 models, released in July 2024, were pretrained on about 15 trillion tokens and took 39.3 million GPU hours. DeepMind’s Chinchilla study found many large models were undertrained, and advised doubling the training tokens whenever model size doubles.
A new model knows nothing, and its settings are far too many to set by hand.
Follow one batch of examples through a training step.
- 1 · predictThe model runs a batch of examples through its current weights and makes a prediction for each one, called the forward pass.
- 2 · scoreA loss function compares each prediction with its label and turns the batch's mistakes into a single number.
- 3 · traceBackpropagation works backwards through the model to find each weight's gradient, which says whether raising or lowering that weight would cut the loss.
- 4 · updateThe optimiser moves every weight a small step against its gradient, with the step size set by the learning rate.
- 5 · repeatThe loop runs over every batch to finish an epoch, then over many epochs until loss on held-out data stops improving.
The update is the learning. Everything before it only works out which way, and how far, each weight should move.
| Who | What they ask | What it works with |
|---|---|---|
| Student running a first tutorial | “Why does my loss jump around instead of going down?” | The learning rate and batch size settings |
| Language model lab | “How many tokens should a model of this size be trained on?” | The compute budget, model size and token count |
| App team with an open model | “Can we adapt a pretrained model to answer our support tickets?” | A few thousand labelled examples from their own tickets |
| Machine learning engineer | “When should this training run stop?” | Validation loss after each epoch |
- Finds working weights from examples, with no one writing the rules by hand.
- The same predict, measure and adjust loop fits a straight line and trains a deep neural network.
- Fine-tuning reuses a pretrained model and needs only hundreds or thousands of task examples.
- Early stopping on validation loss keeps training from running so long that the model overfits.
- A low training loss does not prove the model works on new data; that needs a held-out check.
- No single learning rate suits every problem; each model and dataset has its own best value.
- A long flat stretch of loss can look finished when it is not.
- Training large models is expensive; Llama 3.1 took 39.3 million GPU hours.
Sources used
This explainer is written in original language. The links below support its factual claims.
- docsLinear regression: Gradient descent (Machine Learning Crash Course), Google for Developers · read 27 Sept 2026
- docsLinear regression: Hyperparameters (Machine Learning Crash Course), Google for Developers · read 27 Sept 2026
- docsLinear regression: Loss (Machine Learning Crash Course), Google for Developers · read 27 Sept 2026
- docsMachine Learning Glossary, Google for Developers · read 27 Sept 2026
- docsOptimizing Model Parameters (PyTorch Tutorials), PyTorch · read 27 Sept 2026
- paperLearning representations by back-propagating errors, Nature (Rumelhart, Hinton and Williams, 1986) · read 27 Sept 2026
- paperAdam: A Method for Stochastic Optimization, arXiv (Kingma and Ba; ICLR 2015) · read 27 Sept 2026
- paperLanguage Models are Few-Shot Learners, arXiv (OpenAI, Brown et al.) · read 27 Sept 2026
- paperTraining Compute-Optimal Large Language Models, arXiv (DeepMind, Hoffmann et al.) · read 27 Sept 2026
- repoLlama 3.1 model card, Meta (meta-llama on GitHub) · read 27 Sept 2026
- paperAutomatic differentiation in machine learning: a survey, Baydin, Pearlmutter, Radul, and Siskind (arXiv; JMLR) · read 27 Sept 2026