Concepts

Max tokens

3 min readbeginnerUpdated 28 Sept 2026
1 · In one line

Max tokens sets an upper bound on how many tokens a model may generate for one response.

1 · What it is

Max tokens is an output ceiling. It limits how many tokens a model may generate for one response.

Hugging Face uses max_new_tokens for a limit that ignores the prompt length. vLLM defines max tokens per output sequence. Google documents maxOutputTokens as the maximum number of tokens that can be generated in a response.

OpenAI counts reasoning tokens, visible output tokens and formatting tokens inside one maximum. OpenAI notes that an incomplete response can occur before any visible output tokens are produced.

2 · Why it exists

Generated responses need a clear output boundary.

Runaway lengthvLLM defines a maximum number of tokens per output sequence.
Shared budgetSome reasoning APIs count reasoning tokens and visible output tokens against one maximum.
Library controlHugging Face recommends max new tokens for controlling how many tokens a model generates.
3 · How it works

Count generated tokens until the response ends or the ceiling is reached.

The ceiling limits generated tokens; Hugging Face max new tokens excludes the prompt length.
  1. 1 · setChoose the maximum number of tokens the response may generate.
  2. 2 · generateThe model produces output tokens.
  3. 3 · countThe runtime counts each generated token against the configured ceiling.
  4. 4 · stopThe response becomes incomplete if an OpenAI request reaches max output tokens.

A response can be incomplete when generated tokens reach the ceiling.

4 · Where it's used
WhoWhat they askWhat it works with
API developer“How long may this response become?”The request's output-token ceiling
Product team“Why did this answer stop before its final sentence?”The response status and configured maximum
Inference operator“What is the per-sequence generation cap?”The serving engine's max tokens value
5 · What it solves, and what it doesn't
solves
  • Max tokens places a hard ceiling on generated output.
  • Max new tokens can count generated tokens without counting the prompt.
  • A lower maximum can constrain response length.
doesn't solve
  • A hard output cutoff does not change how Gemini allocates its thinking budget.
  • An OpenAI response is incomplete when generated tokens reach max_output_tokens.
  • OpenAI counts reasoning tokens inside the max_output_tokens limit.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. docsReasoning models, OpenAI · read 28 Sept 2026
  2. docsGenerate content with the Gemini API, Google Cloud · read 28 Sept 2026
  3. docsGeneration, Hugging Face · read 28 Sept 2026
  4. docsSampling parameters, vLLM · read 28 Sept 2026
  5. docsExtended thinking, Anthropic · read 28 Sept 2026