Test-time compute
Test-time compute lets a language model improve its output by using more computation at test time.
Test-time compute lets a language model improve its output by using more computation at test time. The method is given the prompt at test time. Assign an inference-time budget for this prompt.
Repeatedly sample candidate solutions from the model. Self-consistency selects an answer across diverse reasoning paths. Tree of Thoughts can look ahead or backtrack. Search against a process-based verifier reward model.
The best allocation depends on prompt difficulty. In one repeated-sampling study, majority voting and reward models plateaued. OpenAI reported that o1 performance improved with more time spent thinking. Extra test-time compute means more time spent thinking before the answer.
Self-consistency samples diverse reasoning paths instead of taking only the greedy path.
One approach samples diverse reasoning paths and selects the most consistent answer.
- 1 · allocateAssign an inference-time budget for this prompt.
- 2 · exploreRepeatedly sample candidate solutions from the model.
- 3 · scoreSearch against a process-based verifier reward model.
- 4 · answerReturn the selected result.
The method is given the prompt at test time.
| Who | What they ask | What it works with |
|---|---|---|
| Reasoning system | “Which candidate solves this maths problem?” | Multiple sampled solutions |
| Coding agent | “Which proposed program passes the checks?” | Candidate programs and test results |
| Serving team | “How much inference budget should this prompt receive?” | Prompt difficulty and latency budget |
- Repeated sampling can increase the chance that at least one candidate solves a problem.
- Self-consistency selects an answer across diverse reasoning paths.
- Tree of Thoughts can look ahead or backtrack.
- In one repeated-sampling study, majority voting and reward models plateaued.
- The best allocation depends on prompt difficulty.
- Extra test-time compute means more time spent thinking before the answer.
Sources used
This explainer is written in original language. The links below support its factual claims.
- officialLearning to reason with LLMs, OpenAI · read 28 Sept 2026
- paperScaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters, Snell et al. · read 28 Sept 2026
- paperLarge Language Monkeys: Scaling Inference Compute with Repeated Sampling, Brown et al. · read 28 Sept 2026
- paperSelf-Consistency Improves Chain of Thought Reasoning in Language Models, Wang et al. · read 28 Sept 2026
- paperTree of Thoughts: Deliberate Problem Solving with Large Language Models, Yao et al. · read 28 Sept 2026