Concepts

Greedy decoding

3 min readbeginnerUpdated 28 Sept 2026
1 · In one line

Greedy decoding generates a sequence by choosing the most likely next token at every step.

1 · What it is

At each step, greedy decoding selects the token with the highest probability. It adds that token to the generated prefix. The updated prefix enters the next step until generation stops.

The method is fast because it requires one run of the decoder rather than maintaining several prefixes.

On longer text, greedy search can begin to repeat itself. Decoding strategy alone can change generated text even when the language model is unchanged.

2 · Why it exists

A language model scores many possible next tokens, but generation must choose one.

One continuationThe model produces a probability distribution, while the decoder must select the next token.
Serial choicesEach selected token becomes part of the context for the following step.
Simple baselineGreedy decoding requires one run of the underlying decoder rather than maintaining several candidate prefixes.
3 · How it works

Follow one highest-probability choice into the next decoding step.

Greedy decoding commits to the highest-probability token at each step.
  1. 1 · scoreCompute a probability for each possible next token.
  2. 2 · chooseSelect the token with the highest probability.
  3. 3 · appendAdd that token to the generated prefix.
  4. 4 · repeatFeed the updated prefix into the next step until generation stops.

The decision is local: greedy decoding keeps one continuation at each step.

4 · Where it's used
WhoWhat they askWhat it works with
Model engineer“What is the simplest deterministic decoding baseline?”The highest-probability token at each step
Translation researcher“How does greedy output compare with a wider search?”One prefix versus several beam candidates
Product engineer“Why is a long answer repeating itself?”Decoding behavior on longer generations
5 · What it solves, and what it doesn't
solves
  • It turns next-token probabilities into one deterministic continuation.
  • It keeps only one prefix instead of several beam candidates.
  • It provides a simple decoding baseline for comparison.
doesn't solve
  • It selects only the highest-probability token at each step.
  • It can produce repetitive text on longer generations.
  • It does not sample from the probability distribution.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. docsGeneration strategies, Hugging Face · read 28 Sept 2026
  2. docsGeneration, Hugging Face · read 28 Sept 2026
  3. docsNLP From Scratch - Translation with a Sequence to Sequence Network and Attention, PyTorch · read 28 Sept 2026
  4. paperA Stable and Effective Learning Strategy for Trainable Greedy Decoding, Chen, Li, Cho and Bowman · read 28 Sept 2026
  5. paperThe Curious Case of Neural Text Degeneration, Holtzman et al. · read 28 Sept 2026