Concepts

Stop sequences

3 min readbeginnerUpdated 28 Sept 2026
1 · In one line

A stop sequence is a configured character string that stops output generation.

1 · What it is

A stop sequence is a configured character string that stops output generation. Supply one or more character strings that should end generation. Detokenize new token IDs incrementally. Evaluate stop criteria.

On the first match, stop and omit the matched string from the response. Streaming exclusion can require holding back characters from the output. Google accepts a set of up to five character stop sequences.

A maximum-token setting limits the number of generated tokens per output sequence. vLLM excludes matched string stops from returned output. It retains token-ID stops unless they are special tokens. An Anthropic response example includes stop_reason and stop_sequence fields.

2 · Why it exists

Applications often need output to end at a boundary they define.

Text boundaryStop strings terminate generation when the model outputs them.
Stream boundaryvLLM can hold back characters when stop strings are excluded from streamed output.
Different controlsA maximum-token setting limits the number of tokens generated per output sequence.
3 · How it works

Detokenize new token IDs incrementally.

Stop criteria end generation when the configured string is generated.
  1. 1 · configureSupply one or more character strings that should end generation.
  2. 2 · generateDetokenize new token IDs incrementally.
  3. 3 · matchEvaluate stop criteria.
  4. 4 · returnOn the first match, stop and omit the matched string from the response.

A configured stop string terminates generation if the model outputs it.

4 · Where it's used
WhoWhat they askWhat it works with
API developer“Stop before the model begins a second record”A record delimiter
Chat server“End at the assistant-turn marker”The configured turn delimiter
Template runner“Return only the requested field”A following section marker
5 · What it solves, and what it doesn't
solves
  • A stop string terminates generation when it is generated.
  • Google accepts a set of up to five character stop sequences.
  • A server can exclude the matched stop string from returned text.
doesn't solve
  • Streaming exclusion can require holding back characters from the output.
  • A maximum-token setting limits the number of generated tokens per output sequence.
  • vLLM excludes matched string stops from returned output.
  • vLLM retains token-ID stops unless they are special tokens.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. docsGenerating content, Google AI for Developers · read 28 Sept 2026
  2. docsSamplingParams, vLLM · read 28 Sept 2026
  3. docsIncremental detokenizer, vLLM · read 28 Sept 2026
  4. docsGeneration, Hugging Face · read 28 Sept 2026
  5. docsUsing the Messages API, Anthropic · read 28 Sept 2026