Concepts

Beam search

3 min readintermediateUpdated 28 Sept 2026
1 · In one line

Beam search keeps a fixed number of high-scoring partial sequences, expands them, and retains the best candidates at each step.

1 · What it is

Beam search expands each surviving prefix and keeps only the highest-scoring continuations. The number retained is the beam width.

A prefix that is not best now can remain in the beam and lead to a higher-probability sequence later. Beam search is approximate inference. Each pruning step retains only the top B candidates.

Increasing the beam adds decoder work and candidate-management overhead. Standard beam search may produce a list of very similar candidates, and likelihood-maximizing decoding can repeat in open-ended generation. In Transformers, beam-search decoding uses more than one beam with sampling disabled.

2 · Why it exists

The best complete sequence may not begin with the single highest-scoring token.

Early commitmentGreedy decoding keeps one token choice at every step.
Huge searchThe number of possible output sequences grows exponentially with sequence length.
Limited budgetBeam search approximates a wider search by retaining only a fixed number of prefixes.
3 · How it works

Expand two prefixes, score their children, and prune back to a beam of two.

Beam width B controls how many partial sequences survive each pruning step.
  1. 1 · expandGenerate possible next-token extensions for every prefix in the beam.
  2. 2 · scoreScore each extended sequence as a candidate continuation.
  3. 3 · pruneKeep the B highest-scoring prefixes.
  4. 4 · repeatExpand the surviving prefixes at the next step.
  5. 5 · finishReturn the highest-scoring completed sequence under the stopping rule.

A wider beam explores more prefixes but requires more decoder work and candidate management.

4 · Where it's used
WhoWhat they askWhat it works with
Translation researcher“Could a lower-ranked first token lead to a better full translation?”Competing sequence prefixes
Speech engineer“Which transcript has the best accumulated score?”A beam of partial transcriptions
Generation engineer“Why are several returned candidates nearly identical?”Candidate diversity inside the beam
5 · What it solves, and what it doesn't
solves
  • It preserves several sequence prefixes instead of one.
  • It can retain a promising sequence that starts with a lower-ranked token.
  • It provides an approximate search over a very large sequence space.
doesn't solve
  • It does not examine every possible sequence.
  • A larger beam requires more computation and candidate management.
  • Standard beam search can return candidates that differ only slightly.
  • Maximizing sequence likelihood can produce repetitive text in open-ended generation.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. docsGeneration strategies, Hugging Face · read 28 Sept 2026
  2. paperA Stable and Effective Learning Strategy for Trainable Greedy Decoding, Chen, Li, Cho and Bowman · read 28 Sept 2026
  3. paperDiverse Beam Search - Decoding Diverse Solutions from Neural Sequence Models, Vijayakumar et al. · read 28 Sept 2026
  4. paperThe Curious Case of Neural Text Degeneration, Holtzman et al. · read 28 Sept 2026
  5. docsGeneration, Hugging Face · read 28 Sept 2026