Next-token prediction
Next-token prediction estimates which token should follow the tokens already present in a sequence.
During causal language-model training, the next word serves as the label. GPT-1 used a standard language-modeling objective to maximize likelihood. GPT-2 described language modeling by factorizing a sequence probability into conditional probabilities.
The decoder output is converted into predicted next-token probabilities. The next prediction is conditioned on the preceding tokens. PaLM used a decoder-only architecture with a standard left-to-right language-modeling objective. GPT-3 was presented as an autoregressive language model.
A sequence can be learned as a series of conditional prediction problems.
Follow one prefix through a single prediction step.
- 1 · conditionThe model receives the tokens to the left of the next position.
- 2 · scoreThe decoder output is converted into predicted next-token probabilities.
- 3 · compareThe predicted distribution assigns probabilities to possible next tokens.
- 4 · repeatThe next prediction is conditioned on the preceding tokens.
Causal language-model training uses the observed next word as the label.
| Who | What they ask | What it works with |
|---|---|---|
| Model trainer | “Which target should align with this prefix?” | The observed token immediately after the prefix |
| Inference engineer | “Which candidate tokens have the highest scores?” | The next-token probability distribution |
| Product engineer | “Why does generated text arrive one token at a time?” | The repeated autoregressive decoding loop |
| Evaluation team | “How well does the model predict held-out sequences?” | The conditional likelihood of observed tokens |
- Causal language-model training supplies the next word as its label.
- A causal model can assign a probability to each next-token candidate.
- A standard left-to-right objective predicts the next token from preceding tokens.
- A causal model cannot use tokens to the right of the position being predicted.
- GPT-3 text samples can repeat themselves semantically.
- GPT-3 text samples can lose coherence over sufficiently long passages.
Sources used
This explainer is written in original language. The links below support its factual claims.
- docsCausal language modeling, Hugging Face · read 28 Sept 2026
- paperImproving Language Understanding by Generative Pre-Training, OpenAI · read 28 Sept 2026
- paperLanguage Models are Unsupervised Multitask Learners, OpenAI · read 28 Sept 2026
- paperLanguage Models are Few-Shot Learners, Brown and colleagues · read 28 Sept 2026
- paperPaLM: Scaling Language Modeling with Pathways, Chowdhery and colleagues · read 28 Sept 2026
- paperAttention Is All You Need, Vaswani and colleagues · read 28 Sept 2026