Concepts

Top-k sampling

3 min readbeginnerUpdated 28 Sept 2026
1 · In one line

Top-k sampling keeps the k highest-probability next-token candidates, rescales their probabilities and samples one of them.

1 · What it is

Top-k sampling keeps the k highest-probability next-token candidates, rescales their probabilities and samples one of them.

The candidate count is simple and fixed for every decoding step. A constant k can be sub-optimal across varying contexts. If k is small, generated text can become bland or generic. If k is large, the shortlist can include inappropriate candidates.

Fan and colleagues sampled from ten candidates in their story generator. GPT-2 used top-k random sampling with k equal to 40 for generation. vLLM documents zero or minus one as settings that consider all tokens.

2 · Why it exists

Sampling from the full vocabulary can include very unlikely next tokens.

Large vocabularyA language model assigns a probability to every possible next token.
Unlikely tailCompletely random sampling can introduce very unlikely words.
Fixed shortlistTop-k keeps a chosen number of the highest-probability vocabulary tokens.
3 · How it works

Build one fixed-size candidate set.

Top-k keeps k candidates. Their retained probability mass can vary from one step to the next.
  1. 1 · rankIdentify the k most probable next-token candidates.
  2. 2 · keepRetain exactly the k highest-probability tokens.
  3. 3 · rescaleNormalize the retained probabilities over the shortlist.
  4. 4 · sampleRandomly choose one token from that truncated distribution.

Top-k uses a constant k. Under nucleus sampling, the candidate count rises and falls dynamically.

4 · Where it's used
WhoWhat they askWhat it works with
Generation engineer“How many candidates should remain eligible per step?”The top-k value
Runtime engineer“Which setting considers the full vocabulary?”The serving engine's top-k disable value
Evaluation team“Which tokens were removed before sampling?”The ranked probability distribution and k
5 · What it solves, and what it doesn't
solves
  • Top-k excludes all but the k highest-probability candidates.
  • The retained distribution can be sampled rather than always taking rank one.
  • The candidate count is simple and fixed for every decoding step.
doesn't solve
  • A constant k can be sub-optimal across varying contexts.
  • If k is small, generated text can become bland or generic.
  • If k is large, the shortlist can include inappropriate candidates.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. paperThe Curious Case of Neural Text Degeneration, Holtzman et al. · read 28 Sept 2026
  2. docsGeneration, Hugging Face · read 28 Sept 2026
  3. paperHierarchical Neural Story Generation, Fan, Lewis and Dauphin · read 28 Sept 2026
  4. paperLanguage Models are Unsupervised Multitask Learners, OpenAI · read 28 Sept 2026
  5. docsSamplingParams, vLLM · read 28 Sept 2026