Concepts

Loss function

3 min readbeginnerUpdated 28 Sept 2026
1 · In one line

A loss function turns the difference between a model's prediction and its target into a number that training tries to reduce.

1 · What it is

A loss function is a mathematical rule that receives targets and predictions and returns error values. Mean squared error squares each numeric gap. Cross-entropy is built for choosing among classes. Huber loss is quadratic for smaller errors and linear after a threshold.

Training evaluates the loss on examples in a batch and reduces those values, often to an average. Automatic differentiation then finds how that batch loss changes with the model’s parameters. An optimizer uses those gradients to update the parameters.

Different kinds of model call for different loss functions. Loss measured on test examples the model never trained on says more about real quality than the training loss does. Teams also track metrics, whose results are left out of training.

2 · Why it exists

A model needs a precise training signal, not the vague instruction to make better predictions.

Different errorsNumber guesses and class guesses are usually scored with different loss rules.
Unequal costsSquaring an error makes a large regression miss count more strongly than a small one.
Many examplesBy default, Keras averages the loss of every example in a batch into a single value.
3 · How it works

Follow one prediction from target comparison to a batch loss.

A loss function converts prediction errors into one training objective Model predictions and targets enter a highlighted loss function. It emits per-example losses, which are averaged into a batch loss for the optimizer. PREDICTIONS0.70 · 0.20 · 0.90 TARGETS1.00 · 0.00 · 1.00 KEY STEP Loss function compare predictionwith targetrule: (ŷ − y)² Per-example loss.09 · .04 · .01one value each Batch lossmean = .047to optimizer Different task, different ruleclassification → cross-entropy
The loss function defines what training counts as an error and how strongly it counts.
  1. 1 · predictPass a batch of inputs forward through the model to get its outputs.
  2. 2 · compareApply the chosen loss function to each prediction and its target.
  3. 3 · reduceCombine the per-example values, commonly by averaging them across the batch.
  4. 4 · optimizeTake the gradient of the batch loss and nudge the weights so that the loss drops.

The loss is the number training minimizes; a metric can score the model without ever steering the weight updates.

4 · Where it's used
WhoWhat they askWhat it works with
Forecasting team“How costly is this numeric prediction error?”Mean squared or Huber loss
Classifier team“How much probability did the model assign to the correct class?”Cross-entropy loss
Training engineer“Is optimization still reducing the objective?”Training and validation loss curves
5 · What it solves, and what it doesn't
solves
  • It gives training a single number to push down.
  • Mean squared error squares each numeric miss and averages the results.
  • Cross-entropy compares class logits with a class target for classification.
doesn't solve
  • A low training loss does not prove the model will generalize to new examples.
  • Tracking a metric next to the loss does not change training; metric results stay out of the weight updates.
  • A batch average is one number; to see each example's loss you must ask for the unreduced, per-sample values.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. docsMachine Learning Glossary: ML Fundamentals, Google for Developers · read 27 Sept 2026
  2. docsMSELoss, PyTorch · read 27 Sept 2026
  3. docsCrossEntropyLoss, PyTorch · read 27 Sept 2026
  4. docstf.keras.losses.Huber, TensorFlow · read 27 Sept 2026
  5. docsLosses, Keras · read 27 Sept 2026
  6. docsMetrics, Keras · read 27 Sept 2026