Concepts

Test-time compute

3 min readintermediateUpdated 28 Sept 2026
1 · In one line

Test-time compute lets a language model improve its output by using more computation at test time.

1 · What it is

Test-time compute lets a language model improve its output by using more computation at test time. The method is given the prompt at test time. Assign an inference-time budget for this prompt.

Repeatedly sample candidate solutions from the model. Self-consistency selects an answer across diverse reasoning paths. Tree of Thoughts can look ahead or backtrack. Search against a process-based verifier reward model.

The best allocation depends on prompt difficulty. In one repeated-sampling study, majority voting and reward models plateaued. OpenAI reported that o1 performance improved with more time spent thinking. Extra test-time compute means more time spent thinking before the answer.

2 · Why it exists

Self-consistency samples diverse reasoning paths instead of taking only the greedy path.

Single attemptOne sample can miss a solution that another sample finds.
Uneven difficultyThe useful amount and method of computation can depend on the prompt.
Selection gapOpenAI reported reranking 1,000 samples with a learned scoring function.
3 · How it works

One approach samples diverse reasoning paths and selects the most consistent answer.

Self-consistency selects an answer across diverse reasoning paths.
  1. 1 · allocateAssign an inference-time budget for this prompt.
  2. 2 · exploreRepeatedly sample candidate solutions from the model.
  3. 3 · scoreSearch against a process-based verifier reward model.
  4. 4 · answerReturn the selected result.

The method is given the prompt at test time.

4 · Where it's used
WhoWhat they askWhat it works with
Reasoning system“Which candidate solves this maths problem?”Multiple sampled solutions
Coding agent“Which proposed program passes the checks?”Candidate programs and test results
Serving team“How much inference budget should this prompt receive?”Prompt difficulty and latency budget
5 · What it solves, and what it doesn't
solves
  • Repeated sampling can increase the chance that at least one candidate solves a problem.
  • Self-consistency selects an answer across diverse reasoning paths.
  • Tree of Thoughts can look ahead or backtrack.
doesn't solve
  • In one repeated-sampling study, majority voting and reward models plateaued.
  • The best allocation depends on prompt difficulty.
  • Extra test-time compute means more time spent thinking before the answer.