Concepts

Softmax

3 min readbeginnerUpdated 28 Sept 2026
1 · In one line

Softmax determines a probability for each possible class, and those probabilities add up to exactly 1.

1 · What it is

Softmax exponentiates each logit and divides by the sum of all exponentials. Every output uses that same denominator. Higher input scores receive higher output shares.

The result has the same number of elements as the input, but every value is between 0 and 1 and the values sum to 1. In the figure, the illustrative logits 1, 2 and 0 produce softmax values of approximately 0.24, 0.67 and 0.09.

Transformer attention applies softmax to scaled query-key products to obtain weights on values.

2 · Why it exists

Softmax maps a model's class scores to probabilities.

Scores need mappingA multi-class model outputs one value for each possible class.
Classes competeSoftmax returns an output vector with the same number of elements as its input vector.
Scale changes certaintyMultiplying all logits by a large positive number makes the largest softmax share approach 1.
3 · How it works

Follow three logits into one distribution.

Softmax exponentiates every score, then divides every result by the same total.
  1. 1 · scoreThe model produces one real-valued logit for each possible class.
  2. 2 · exponentiateSoftmax applies the exponential function to each logit, making every transformed value positive.
  3. 3 · totalIt adds those exponentials to form one shared normalizing denominator.
  4. 4 · divideEach exponential is divided by the shared total, so the outputs lie between 0 and 1 and sum to 1.

Every output uses the same denominator: the sum of all exponentials.

4 · Where it's used
WhoWhat they askWhat it works with
Image classifier“Which object class best matches this image?”One logit per candidate class
Recommendation team“Which item is most likely to be selected?”Scores for all candidate items in the softmax set
Language model“Which token should come next?”Vocabulary logits for the next position
Attention layer“How much weight should each value receive?”Scaled query-key compatibility scores
5 · What it solves, and what it doesn't
solves
  • Its output has the same shape as its input.
  • The output values sum to 1, so they can represent one probability distribution over mutually competing choices.
  • Higher input scores receive higher output shares.
doesn't solve
  • It does not make class predictions independent; changing one logit changes the shared denominator for every class.
  • Full softmax can be expensive when the number of classes is very large.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. docsMachine Learning Glossary - softmax, Google for Developers · read 27 Sept 2026
  2. docsSoftmax, PyTorch · read 27 Sept 2026
  3. docsSoftmax DNN for recommendation, Google for Developers · read 27 Sept 2026
  4. docsNeural networks: Multi-class classification, Google for Developers · read 27 Sept 2026
  5. paperAttention Is All You Need, Vaswani et al. · read 27 Sept 2026