Max tokens
Max tokens sets an upper bound on how many tokens a model may generate for one response.
Max tokens is an output ceiling. It limits how many tokens a model may generate for one response.
Hugging Face uses max_new_tokens for a limit that ignores the prompt length. vLLM defines max tokens per output sequence. Google documents maxOutputTokens as the maximum number of tokens that can be generated in a response.
OpenAI counts reasoning tokens, visible output tokens and formatting tokens inside one maximum. OpenAI notes that an incomplete response can occur before any visible output tokens are produced.
Generated responses need a clear output boundary.
Count generated tokens until the response ends or the ceiling is reached.
- 1 · setChoose the maximum number of tokens the response may generate.
- 2 · generateThe model produces output tokens.
- 3 · countThe runtime counts each generated token against the configured ceiling.
- 4 · stopThe response becomes incomplete if an OpenAI request reaches max output tokens.
A response can be incomplete when generated tokens reach the ceiling.
| Who | What they ask | What it works with |
|---|---|---|
| API developer | “How long may this response become?” | The request's output-token ceiling |
| Product team | “Why did this answer stop before its final sentence?” | The response status and configured maximum |
| Inference operator | “What is the per-sequence generation cap?” | The serving engine's max tokens value |
- Max tokens places a hard ceiling on generated output.
- Max new tokens can count generated tokens without counting the prompt.
- A lower maximum can constrain response length.
- A hard output cutoff does not change how Gemini allocates its thinking budget.
- An OpenAI response is incomplete when generated tokens reach max_output_tokens.
- OpenAI counts reasoning tokens inside the max_output_tokens limit.
Sources used
This explainer is written in original language. The links below support its factual claims.
- docsReasoning models, OpenAI · read 28 Sept 2026
- docsGenerate content with the Gemini API, Google Cloud · read 28 Sept 2026
- docsGeneration, Hugging Face · read 28 Sept 2026
- docsSampling parameters, vLLM · read 28 Sept 2026
- docsExtended thinking, Anthropic · read 28 Sept 2026