Speculative decoding
Speculative decoding uses a target model to evaluate guesses in parallel.
A faster approximation proposes a short continuation. The target model evaluates the guesses in parallel. If the target rejects drafted tokens, the rest of that proposal batch is discarded. The algorithm samples a correction for the first rejected token.
The exact sampling method preserves the target model’s distribution. Medusa adds decoding heads that predict multiple subsequent tokens in parallel, while SpecInfer organizes proposals as a token tree.
Staged speculative decoding uses a fast draft model to anticipate the target and batch queries to it. A rejected guess wastes computation; benefits can also shrink at larger serving batch sizes.
Autoregressive decoding normally produces one dependent token at a time.
Draft a short continuation, verify it once, then commit only the accepted prefix.
- 1 · draftUse a faster approximation model to propose a short token continuation.
- 2 · verifyEvaluate the proposed continuation with the target model in parallel.
- 3 · acceptIf the target rejects drafted tokens, discard the rest of that proposal batch.
- 4 · correctSample a correction for the first rejection, then begin another round.
The exact sampling algorithm preserves the target model's output distribution.
| Who | What they ask | What it works with |
|---|---|---|
| Inference engineer | “Can one target-model call advance generation by several tokens?” | Accepted tokens per verification round |
| Runtime developer | “Does a faster draft model reduce small-batch decoding latency?” | End-to-end speculative speedup |
| Systems researcher | “How do draft length and acceptance affect wasted work?” | Acceptance rate and draft cost |
- It can generate multiple tokens from one target-model verification call.
- Rejection sampling can preserve the target model's output distribution.
- SpecInfer organizes speculated tokens in a token tree.
- A rejected guess wastes computation.
- Speedup can plateau or degrade when the proposed continuation becomes too long.
- Benefits can shrink as serving batch size increases.
Sources used
This explainer is written in original language. The links below support its factual claims.
- paperFast Inference from Transformers via Speculative Decoding, ICML / arXiv · read 28 Sept 2026
- paperAccelerating Large Language Model Decoding with Speculative Sampling, arXiv · read 28 Sept 2026
- paperSpecInfer - Accelerating Generative Large Language Model Serving with Tree-based Speculative Inference and Verification, ASPLOS / arXiv · read 28 Sept 2026
- paperMedusa - Simple LLM Inference Acceleration Framework with Multiple Decoding Heads, ICML / arXiv · read 28 Sept 2026
- paperAccelerating LLM Inference with Staged Speculative Decoding, ICML / arXiv · read 28 Sept 2026