Concepts

Logits

3 min readbeginnerUpdated 28 Sept 2026
1 · In one line

Logits are a model's raw, unnormalized prediction scores.

1 · What it is

Logits are a model’s raw, unnormalized prediction scores. The model produces one raw score for each candidate class.

Softmax converts the vector into one probability per class. A cross-entropy loss can act directly on the logits. PyTorch cross-entropy accepts unnormalized logits for each class.

The Transformer used a learned linear transformation and softmax to produce predicted next-token probabilities.

2 · Why it exists

A classifier produces raw scores before softmax forms a probability distribution.

Raw outputGoogle defines logits as a vector of raw, non-normalized predictions.
One per classA softmax output layer converts the logit computed for each class into a probability.
Loss inputPyTorch cross-entropy accepts unnormalized logits for each class.
3 · How it works

Follow three class logits into softmax.

Logits are raw scores; softmax normalizes the full vector into probabilities.
  1. 1 · scoreThe model produces one raw score for each candidate class.
  2. 2 · collectThose scores form an unnormalized logit vector.
  3. 3 · normalizeSoftmax converts the vector into one probability per class.
  4. 4 · useA cross-entropy loss can act directly on the logits.

A logit is not itself a probability; the vector has not yet been normalized.

4 · Where it's used
WhoWhat they askWhat it works with
Classifier engineer“Which raw class score did the network produce?”The output logit vector
Language-model engineer“Which tokens have the highest pre-softmax scores?”Next-token logits
Training engineer“Should cross-entropy receive logits or probabilities?”The loss function's input contract
5 · What it solves, and what it doesn't
solves
  • Logits preserve the model's raw class scores before normalization.
  • Softmax can turn a logit vector into normalized class probabilities.
  • Cross-entropy implementations can consume logits directly.
doesn't solve
  • A single logit does not give a normalized probability by itself.
  • Logits ordinarily still pass through a normalization function such as softmax.
  • Some cross-entropy losses expect unnormalized logits rather than probabilities.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. docsMachine Learning Glossary, Google for Developers · read 28 Sept 2026
  2. paperDistilling the Knowledge in a Neural Network, Hinton, Vinyals and Dean · read 28 Sept 2026
  3. paperAttention Is All You Need, Vaswani et al. · read 28 Sept 2026
  4. docstf.keras.ops.categorical_crossentropy, TensorFlow · read 28 Sept 2026
  5. docsCrossEntropyLoss, PyTorch · read 28 Sept 2026