Concepts

Top-p sampling

3 min readintermediateUpdated 28 Sept 2026
1 · In one line

Top-p keeps the smallest set of most probable tokens whose probabilities add up to at least p.

1 · What it is

Top-p sampling is also called nucleus sampling. Select highest-probability tokens while tracking their cumulative probability mass. Retain the smallest prefix that reaches or exceeds the threshold.

Rescale the retained probabilities. Nucleus sampling samples from the top-p portion of probability mass. Top-p lets the number of eligible tokens rise and fall dynamically.

Holtzman and colleagues introduced nucleus sampling to truncate the unreliable probability tail. The authors’ repository directs users to a Hugging Face implementation of nucleus sampling. vLLM uses one to consider all tokens.

2 · Why it exists

Top-k sampling can fail for any one choice of k.

Confidence changesThe useful candidate pool can expand and contract from one decoding step to the next.
Long tailNucleus sampling truncates the unreliable tail of the probability distribution.
Adaptive setThe number of candidates rises and falls dynamically.
3 · How it works

Build one nucleus from a sorted distribution.

Top-p chooses probability mass, so the number of surviving tokens can change at every step.
  1. 1 · sortSelect highest-probability tokens while tracking their cumulative probability mass.
  2. 2 · accumulateSelect highest-probability tokens whose cumulative probability mass exceeds p.
  3. 3 · keepRetain the smallest prefix that reaches or exceeds the threshold.
  4. 4 · sampleRescale the retained probabilities.

The threshold p controls the cumulative probability of the top tokens considered.

4 · Where it's used
WhoWhat they askWhat it works with
Generation engineer“How can the candidate count expand and contract dynamically?”The top-p threshold
Evaluation team“Which candidates were eligible at this decoding step?”The cumulative probability nucleus
Runtime engineer“Which setting disables top-p filtering?”The serving engine's sampling parameters
5 · What it solves, and what it doesn't
solves
  • Top-p adapts the number of eligible tokens to the distribution at each step.
  • It removes the probability tail outside the selected nucleus.
  • The nucleus always contains probability mass of at least p.
doesn't solve
  • Setting p too low can over-truncate the distribution and resemble greedy decoding.
  • Nucleus sampling is a stochastic decoding method.
  • A p value of one disables top-p filtering by considering all tokens.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. paperThe Curious Case of Neural Text Degeneration, Holtzman et al. · read 28 Sept 2026
  2. docsGeneration, Hugging Face · read 28 Sept 2026
  3. docsSamplingParams, vLLM · read 28 Sept 2026
  4. paperConformal Nucleus Sampling, Ravfogel, Goldberg and Goldberger · read 28 Sept 2026
  5. repoNucleus sampling and generations, Holtzman et al. · read 28 Sept 2026