Beam search
Beam search keeps a fixed number of high-scoring partial sequences, expands them, and retains the best candidates at each step.
Beam search expands each surviving prefix and keeps only the highest-scoring continuations. The number retained is the beam width.
A prefix that is not best now can remain in the beam and lead to a higher-probability sequence later. Beam search is approximate inference. Each pruning step retains only the top B candidates.
Increasing the beam adds decoder work and candidate-management overhead. Standard beam search may produce a list of very similar candidates, and likelihood-maximizing decoding can repeat in open-ended generation. In Transformers, beam-search decoding uses more than one beam with sampling disabled.
The best complete sequence may not begin with the single highest-scoring token.
Expand two prefixes, score their children, and prune back to a beam of two.
- 1 · expandGenerate possible next-token extensions for every prefix in the beam.
- 2 · scoreScore each extended sequence as a candidate continuation.
- 3 · pruneKeep the B highest-scoring prefixes.
- 4 · repeatExpand the surviving prefixes at the next step.
- 5 · finishReturn the highest-scoring completed sequence under the stopping rule.
A wider beam explores more prefixes but requires more decoder work and candidate management.
| Who | What they ask | What it works with |
|---|---|---|
| Translation researcher | “Could a lower-ranked first token lead to a better full translation?” | Competing sequence prefixes |
| Speech engineer | “Which transcript has the best accumulated score?” | A beam of partial transcriptions |
| Generation engineer | “Why are several returned candidates nearly identical?” | Candidate diversity inside the beam |
- It preserves several sequence prefixes instead of one.
- It can retain a promising sequence that starts with a lower-ranked token.
- It provides an approximate search over a very large sequence space.
- It does not examine every possible sequence.
- A larger beam requires more computation and candidate management.
- Standard beam search can return candidates that differ only slightly.
- Maximizing sequence likelihood can produce repetitive text in open-ended generation.
Sources used
This explainer is written in original language. The links below support its factual claims.
- docsGeneration strategies, Hugging Face · read 28 Sept 2026
- paperA Stable and Effective Learning Strategy for Trainable Greedy Decoding, Chen, Li, Cho and Bowman · read 28 Sept 2026
- paperDiverse Beam Search - Decoding Diverse Solutions from Neural Sequence Models, Vijayakumar et al. · read 28 Sept 2026
- paperThe Curious Case of Neural Text Degeneration, Holtzman et al. · read 28 Sept 2026
- docsGeneration, Hugging Face · read 28 Sept 2026