Chain-of-thought prompting
Chain-of-thought prompting asks for or demonstrates intermediate reasoning text before a final answer, but that text is not a guaranteed view of hidden model internals.
Chain-of-thought prompting puts intermediate reasoning into generated text. The original pattern showed a model several worked solutions, each with steps before the answer, then asked a new question in the same format. Later work showed that a short step-by-step cue could also elicit intermediate text without examples.
This technique concerns visible tokens: words, equations, or subgoals emitted before a final answer. It should not be confused with neural activations or a provider’s private reasoning trace. Some hosted reasoning models spend tokens on an internal process that the provider does not expose. An API may return a summary of that reasoning, but the summary is still not the raw trace.
There is another limit: a convincing rationale may not faithfully identify what caused the prediction. In one study, researchers nudged models by reordering multiple-choice options so the expected letter was always (A). The nudge pulled the answers, yet the written explanations routinely left it out. Treat reasoning text as a working surface. Check calculations, run code, inspect sources, or use a deterministic verifier when correctness matters.
Many problems break into smaller steps, while a bare answer hides where an error entered.
Separate visible reasoning text from private model computation.
- 1 · showThe prompt can include worked examples whose answers contain intermediate reasoning steps.
- 2 · askA new multistep problem follows the same input-and-answer pattern.
- 3 · reasonThe model generates intermediate text before committing to the final answer.
- 4 · verifyA person or program checks the answer and any externally verifiable steps instead of treating the prose as proof.
Visible chain-of-thought is text the model produced; private reasoning tokens and neural activations are different things.
| Who | What they ask | What it works with |
|---|---|---|
| Math tutor | “Where did this solution go wrong?” | Generated steps and independently checked calculations |
| Analyst | “Which constraints rule out each option?” | A concise justification paired with the final choice |
| Developer | “Why does this test case fail?” | An explicit, testable decomposition of the problem |
| Evaluator | “Did the explanation use the evidence it claims?” | Counterfactual tests of the generated rationale |
- Few-shot chain-of-thought demonstrations improved results in the original paper on arithmetic, commonsense, and symbolic reasoning tasks.
- Intermediate text can make individual calculations or assumptions easier to inspect.
- A reasoning summary can outline how a response was reached without exposing the raw reasoning tokens.
- A fluent explanation is not proof that it faithfully reports the cause of the model's answer.
- Providers may keep raw reasoning traces hidden and expose only a summary or final answer.
- Reasoning text does not make a wrong premise, missing fact, or unchecked calculation correct.
Sources used
This explainer is written in original language. The links below support its factual claims.
- paperChain-of-Thought Prompting Elicits Reasoning in Large Language Models, Wei et al. · read 27 Sept 2026
- paperLarge Language Models are Zero-Shot Reasoners, Kojima et al. · read 27 Sept 2026
- officialLearning to reason with LLMs, OpenAI · read 27 Sept 2026
- docsReasoning models, OpenAI · read 27 Sept 2026
- paperLanguage Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting, Turpin et al. · read 27 Sept 2026