Concepts

Speculative decoding

4 min readadvancedUpdated 28 Sept 2026
1 · In one line

Speculative decoding uses a target model to evaluate guesses in parallel.

1 · What it is

A faster approximation proposes a short continuation. The target model evaluates the guesses in parallel. If the target rejects drafted tokens, the rest of that proposal batch is discarded. The algorithm samples a correction for the first rejected token.

The exact sampling method preserves the target model’s distribution. Medusa adds decoding heads that predict multiple subsequent tokens in parallel, while SpecInfer organizes proposals as a token tree.

Staged speculative decoding uses a fast draft model to anticipate the target and batch queries to it. A rejected guess wastes computation; benefits can also shrink at larger serving batch sizes.

2 · Why it exists

Autoregressive decoding normally produces one dependent token at a time.

Sequential stepsA selected token is fed back to produce logits for the following token.
Low utilizationSmall-batch decoding can have low arithmetic intensity.
Parallel checkShort draft continuations can be scored by the target model in parallel.
3 · How it works

Draft a short continuation, verify it once, then commit only the accepted prefix.

The target model evaluates drafted guesses in parallel; the token example is illustrative.
  1. 1 · draftUse a faster approximation model to propose a short token continuation.
  2. 2 · verifyEvaluate the proposed continuation with the target model in parallel.
  3. 3 · acceptIf the target rejects drafted tokens, discard the rest of that proposal batch.
  4. 4 · correctSample a correction for the first rejection, then begin another round.

The exact sampling algorithm preserves the target model's output distribution.

4 · Where it's used
WhoWhat they askWhat it works with
Inference engineer“Can one target-model call advance generation by several tokens?”Accepted tokens per verification round
Runtime developer“Does a faster draft model reduce small-batch decoding latency?”End-to-end speculative speedup
Systems researcher“How do draft length and acceptance affect wasted work?”Acceptance rate and draft cost
5 · What it solves, and what it doesn't
solves
  • It can generate multiple tokens from one target-model verification call.
  • Rejection sampling can preserve the target model's output distribution.
  • SpecInfer organizes speculated tokens in a token tree.
doesn't solve
  • A rejected guess wastes computation.
  • Speedup can plateau or degrade when the proposed continuation becomes too long.
  • Benefits can shrink as serving batch size increases.