Concepts
Greedy decoding
1 · In one line
Greedy decoding generates a sequence by choosing the most likely next token at every step.
1 · What it is
At each step, greedy decoding selects the token with the highest probability. It adds that token to the generated prefix. The updated prefix enters the next step until generation stops.
The method is fast because it requires one run of the decoder rather than maintaining several prefixes.
On longer text, greedy search can begin to repeat itself. Decoding strategy alone can change generated text even when the language model is unchanged.
A language model scores many possible next tokens, but generation must choose one.
One continuationThe model produces a probability distribution, while the decoder must select the next token.
Serial choicesEach selected token becomes part of the context for the following step.
Simple baselineGreedy decoding requires one run of the underlying decoder rather than maintaining several candidate prefixes.
Follow one highest-probability choice into the next decoding step.
- 1 · scoreCompute a probability for each possible next token.
- 2 · chooseSelect the token with the highest probability.
- 3 · appendAdd that token to the generated prefix.
- 4 · repeatFeed the updated prefix into the next step until generation stops.
The decision is local: greedy decoding keeps one continuation at each step.
| Who | What they ask | What it works with |
|---|---|---|
| Model engineer | “What is the simplest deterministic decoding baseline?” | The highest-probability token at each step |
| Translation researcher | “How does greedy output compare with a wider search?” | One prefix versus several beam candidates |
| Product engineer | “Why is a long answer repeating itself?” | Decoding behavior on longer generations |
solves
- It turns next-token probabilities into one deterministic continuation.
- It keeps only one prefix instead of several beam candidates.
- It provides a simple decoding baseline for comparison.
doesn't solve
- It selects only the highest-probability token at each step.
- It can produce repetitive text on longer generations.
- It does not sample from the probability distribution.
6 · Go deeper
Sources used
This explainer is written in original language. The links below support its factual claims.
- docsGeneration strategies, Hugging Face · read 28 Sept 2026
- docsGeneration, Hugging Face · read 28 Sept 2026
- docsNLP From Scratch - Translation with a Sequence to Sequence Network and Attention, PyTorch · read 28 Sept 2026
- paperA Stable and Effective Learning Strategy for Trainable Greedy Decoding, Chen, Li, Cho and Bowman · read 28 Sept 2026
- paperThe Curious Case of Neural Text Degeneration, Holtzman et al. · read 28 Sept 2026