Concepts

Reasoning models

4 min readintermediateUpdated 28 Sept 2026
1 · In one line

Reasoning models use intermediate reasoning tokens before producing a final response.

1 · What it is

Reasoning models generate intermediate reasoning tokens before they return a final response. Those tokens let the model break down a prompt, inspect alternatives and work through multiple steps.

OpenAI records reasoning-token counts in output-token details but does not expose the tokens through the API. DeepSeek returns reasoning_content separately from the final content. Google can return thought summaries. Anthropic supports a thinking-token budget on supported models.

Those tokens occupy context-window space and are billed as output tokens in OpenAI’s API. If the overall ceiling is too low, an API can end the request before a visible answer appears.

2 · Why it exists

Some tasks benefit from planning before the final response.

Multiple stepsHarder tasks may require the model to break down a prompt and consider more than one approach.
Tool workReasoning models can think between tool calls in supported OpenAI models.
Verifiable tasksDeepSeek-R1-Zero relied exclusively on reinforcement learning without supervised fine-tuning.
3 · How it works

Separate intermediate reasoning from the final answer.

OpenAI says reasoning tokens are used before a response. Google applies max_output_tokens to the combined total of thinking tokens and output tokens.
  1. 1 · readThe model receives the prompt.
  2. 2 · reasonIt spends reasoning tokens breaking down the prompt and considering approaches.
  3. 3 · actSupported models may reason between tool calls.
  4. 4 · answerThe model produces a separate final response.

Google returns only the final output by default.

4 · Where it's used
WhoWhat they askWhat it works with
Software engineer“Which model should handle a multi-step debugging task?”Reasoning support and effort controls
Research team“How many reasoning tokens did this request use?”Output token details
Agent developer“Can the model think between tool calls?”Interleaved thinking support
5 · What it solves, and what it doesn't
solves
  • Reasoning tokens give a model room to break down a prompt before answering.
  • Reasoning models can inspect alternatives during generation.
  • Some APIs let developers control reasoning effort or a thinking budget.
doesn't solve
  • Reasoning tokens still occupy context space and count as output tokens in OpenAI's API.
  • A low output ceiling can end a request before any visible answer appears.
  • Simple requests may need only minimal or low thinking.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. docsReasoning models, OpenAI · read 28 Sept 2026
  2. docsExtended thinking, Anthropic · read 28 Sept 2026
  3. docsGemini thinking, Google AI for Developers · read 28 Sept 2026
  4. paperDeepSeek-R1, DeepSeek-AI · read 28 Sept 2026
  5. docsThinking mode, DeepSeek · read 28 Sept 2026