Sampling temperature
Temperature enters softmax by dividing each logit by T.
Temperature shapes a probability distribution during sampling. The formula divides each class logit by the same temperature T.
Temperature can skew probability toward high-probability events. Higher temperature produces a softer probability distribution over classes. Temperature T is normally set to one.
Temperature alone does not sufficiently suppress an unreliable probability tail. Lower temperature can reduce diversity. vLLM documents lower values as more deterministic and higher values as more random. Hugging Face documents temperature as a value that modulates next-token probabilities. Fan and colleagues tuned a softmax temperature at generation time.
Temperature shapes the next-token probability distribution.
Divide one logit vector by temperature before softmax.
- 1 · receiveStart with the model's next-token logits.
- 2 · divideDivide every logit by temperature T.
- 3 · normalizeApply softmax to the scaled logits.
- 4 · sampleUse the resulting distribution for sampling.
The distillation paper says temperature T is normally set to 1.
| Who | What they ask | What it works with |
|---|---|---|
| Generation engineer | “Should the next-token distribution be sharper or flatter?” | The sampling temperature |
| Evaluation team | “Which temperature was used for this generation run?” | The decoding configuration |
| Distillation researcher | “How can class probabilities form softer targets?” | A higher softmax temperature |
- Higher temperature produces a softer probability distribution over classes.
- The formula divides each class logit by the same temperature T.
- Temperature T is normally set to one.
- vLLM documents lower values as more deterministic and higher values as more random.
- Lower temperature can reduce diversity.
- Temperature alone does not sufficiently suppress an unreliable probability tail.
- In the paper's examples, sampling with temperature can still produce incoherent text.
Sources used
This explainer is written in original language. The links below support its factual claims.
- paperDistilling the Knowledge in a Neural Network, Hinton, Vinyals and Dean · read 28 Sept 2026
- paperThe Curious Case of Neural Text Degeneration, Holtzman et al. · read 28 Sept 2026
- docsGeneration, Hugging Face · read 28 Sept 2026
- docsSamplingParams, vLLM · read 28 Sept 2026
- paperHierarchical Neural Story Generation, Fan, Lewis and Dauphin · read 28 Sept 2026