Concepts
Logits
1 · In one line
Logits are a model's raw, unnormalized prediction scores.
1 · What it is
Logits are a model’s raw, unnormalized prediction scores. The model produces one raw score for each candidate class.
Softmax converts the vector into one probability per class. A cross-entropy loss can act directly on the logits. PyTorch cross-entropy accepts unnormalized logits for each class.
The Transformer used a learned linear transformation and softmax to produce predicted next-token probabilities.
A classifier produces raw scores before softmax forms a probability distribution.
Raw outputGoogle defines logits as a vector of raw, non-normalized predictions.
One per classA softmax output layer converts the logit computed for each class into a probability.
Loss inputPyTorch cross-entropy accepts unnormalized logits for each class.
Follow three class logits into softmax.
- 1 · scoreThe model produces one raw score for each candidate class.
- 2 · collectThose scores form an unnormalized logit vector.
- 3 · normalizeSoftmax converts the vector into one probability per class.
- 4 · useA cross-entropy loss can act directly on the logits.
A logit is not itself a probability; the vector has not yet been normalized.
| Who | What they ask | What it works with |
|---|---|---|
| Classifier engineer | “Which raw class score did the network produce?” | The output logit vector |
| Language-model engineer | “Which tokens have the highest pre-softmax scores?” | Next-token logits |
| Training engineer | “Should cross-entropy receive logits or probabilities?” | The loss function's input contract |
solves
- Logits preserve the model's raw class scores before normalization.
- Softmax can turn a logit vector into normalized class probabilities.
- Cross-entropy implementations can consume logits directly.
doesn't solve
- A single logit does not give a normalized probability by itself.
- Logits ordinarily still pass through a normalization function such as softmax.
- Some cross-entropy losses expect unnormalized logits rather than probabilities.
6 · Go deeper
Sources used
This explainer is written in original language. The links below support its factual claims.
- docsMachine Learning Glossary, Google for Developers · read 28 Sept 2026
- paperDistilling the Knowledge in a Neural Network, Hinton, Vinyals and Dean · read 28 Sept 2026
- paperAttention Is All You Need, Vaswani et al. · read 28 Sept 2026
- docstf.keras.ops.categorical_crossentropy, TensorFlow · read 28 Sept 2026
- docsCrossEntropyLoss, PyTorch · read 28 Sept 2026