TemperatureConcepts

Sampling temperature

3 min readintermediateUpdated 28 Sept 2026
1 · In one line

Temperature enters softmax by dividing each logit by T.

1 · What it is

Temperature shapes a probability distribution during sampling. The formula divides each class logit by the same temperature T.

Temperature can skew probability toward high-probability events. Higher temperature produces a softer probability distribution over classes. Temperature T is normally set to one.

Temperature alone does not sufficiently suppress an unreliable probability tail. Lower temperature can reduce diversity. vLLM documents lower values as more deterministic and higher values as more random. Hugging Face documents temperature as a value that modulates next-token probabilities. Fan and colleagues tuned a softmax temperature at generation time.

2 · Why it exists

Temperature shapes the next-token probability distribution.

Same logitsTemperature shapes a probability distribution during sampling.
Sharper choicesTemperature can skew probability toward high-probability events.
Softer choicesHigher temperature produces a softer probability distribution over classes.
3 · How it works

Divide one logit vector by temperature before softmax.

Softmax uses each logit divided by temperature T.
  1. 1 · receiveStart with the model's next-token logits.
  2. 2 · divideDivide every logit by temperature T.
  3. 3 · normalizeApply softmax to the scaled logits.
  4. 4 · sampleUse the resulting distribution for sampling.

The distillation paper says temperature T is normally set to 1.

4 · Where it's used
WhoWhat they askWhat it works with
Generation engineer“Should the next-token distribution be sharper or flatter?”The sampling temperature
Evaluation team“Which temperature was used for this generation run?”The decoding configuration
Distillation researcher“How can class probabilities form softer targets?”A higher softmax temperature
5 · What it solves, and what it doesn't
solves
  • Higher temperature produces a softer probability distribution over classes.
  • The formula divides each class logit by the same temperature T.
  • Temperature T is normally set to one.
  • vLLM documents lower values as more deterministic and higher values as more random.
doesn't solve
  • Lower temperature can reduce diversity.
  • Temperature alone does not sufficiently suppress an unreliable probability tail.
  • In the paper's examples, sampling with temperature can still produce incoherent text.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. paperDistilling the Knowledge in a Neural Network, Hinton, Vinyals and Dean · read 28 Sept 2026
  2. paperThe Curious Case of Neural Text Degeneration, Holtzman et al. · read 28 Sept 2026
  3. docsGeneration, Hugging Face · read 28 Sept 2026
  4. docsSamplingParams, vLLM · read 28 Sept 2026
  5. paperHierarchical Neural Story Generation, Fan, Lewis and Dauphin · read 28 Sept 2026