Softmax
Softmax determines a probability for each possible class, and those probabilities add up to exactly 1.
Softmax exponentiates each logit and divides by the sum of all exponentials. Every output uses that same denominator. Higher input scores receive higher output shares.
The result has the same number of elements as the input, but every value is between 0 and 1 and the values sum to 1. In the figure, the illustrative logits 1, 2 and 0 produce softmax values of approximately 0.24, 0.67 and 0.09.
Transformer attention applies softmax to scaled query-key products to obtain weights on values.
Softmax maps a model's class scores to probabilities.
Follow three logits into one distribution.
- 1 · scoreThe model produces one real-valued logit for each possible class.
- 2 · exponentiateSoftmax applies the exponential function to each logit, making every transformed value positive.
- 3 · totalIt adds those exponentials to form one shared normalizing denominator.
- 4 · divideEach exponential is divided by the shared total, so the outputs lie between 0 and 1 and sum to 1.
Every output uses the same denominator: the sum of all exponentials.
| Who | What they ask | What it works with |
|---|---|---|
| Image classifier | “Which object class best matches this image?” | One logit per candidate class |
| Recommendation team | “Which item is most likely to be selected?” | Scores for all candidate items in the softmax set |
| Language model | “Which token should come next?” | Vocabulary logits for the next position |
| Attention layer | “How much weight should each value receive?” | Scaled query-key compatibility scores |
- Its output has the same shape as its input.
- The output values sum to 1, so they can represent one probability distribution over mutually competing choices.
- Higher input scores receive higher output shares.
- It does not make class predictions independent; changing one logit changes the shared denominator for every class.
- Full softmax can be expensive when the number of classes is very large.
Sources used
This explainer is written in original language. The links below support its factual claims.
- docsMachine Learning Glossary - softmax, Google for Developers · read 27 Sept 2026
- docsSoftmax, PyTorch · read 27 Sept 2026
- docsSoftmax DNN for recommendation, Google for Developers · read 27 Sept 2026
- docsNeural networks: Multi-class classification, Google for Developers · read 27 Sept 2026
- paperAttention Is All You Need, Vaswani et al. · read 27 Sept 2026