Loss function
A loss function turns the difference between a model's prediction and its target into a number that training tries to reduce.
A loss function is a mathematical rule that receives targets and predictions and returns error values. Mean squared error squares each numeric gap. Cross-entropy is built for choosing among classes. Huber loss is quadratic for smaller errors and linear after a threshold.
Training evaluates the loss on examples in a batch and reduces those values, often to an average. Automatic differentiation then finds how that batch loss changes with the model’s parameters. An optimizer uses those gradients to update the parameters.
Different kinds of model call for different loss functions. Loss measured on test examples the model never trained on says more about real quality than the training loss does. Teams also track metrics, whose results are left out of training.
A model needs a precise training signal, not the vague instruction to make better predictions.
Follow one prediction from target comparison to a batch loss.
- 1 · predictPass a batch of inputs forward through the model to get its outputs.
- 2 · compareApply the chosen loss function to each prediction and its target.
- 3 · reduceCombine the per-example values, commonly by averaging them across the batch.
- 4 · optimizeTake the gradient of the batch loss and nudge the weights so that the loss drops.
The loss is the number training minimizes; a metric can score the model without ever steering the weight updates.
| Who | What they ask | What it works with |
|---|---|---|
| Forecasting team | “How costly is this numeric prediction error?” | Mean squared or Huber loss |
| Classifier team | “How much probability did the model assign to the correct class?” | Cross-entropy loss |
| Training engineer | “Is optimization still reducing the objective?” | Training and validation loss curves |
- It gives training a single number to push down.
- Mean squared error squares each numeric miss and averages the results.
- Cross-entropy compares class logits with a class target for classification.
- A low training loss does not prove the model will generalize to new examples.
- Tracking a metric next to the loss does not change training; metric results stay out of the weight updates.
- A batch average is one number; to see each example's loss you must ask for the unreduced, per-sample values.
Sources used
This explainer is written in original language. The links below support its factual claims.
- docsMachine Learning Glossary: ML Fundamentals, Google for Developers · read 27 Sept 2026
- docsMSELoss, PyTorch · read 27 Sept 2026
- docsCrossEntropyLoss, PyTorch · read 27 Sept 2026
- docstf.keras.losses.Huber, TensorFlow · read 27 Sept 2026
- docsLosses, Keras · read 27 Sept 2026
- docsMetrics, Keras · read 27 Sept 2026