Concepts

Next-token prediction

3 min readbeginnerUpdated 28 Sept 2026
1 · In one line

Next-token prediction estimates which token should follow the tokens already present in a sequence.

1 · What it is

During causal language-model training, the next word serves as the label. GPT-1 used a standard language-modeling objective to maximize likelihood. GPT-2 described language modeling by factorizing a sequence probability into conditional probabilities.

The decoder output is converted into predicted next-token probabilities. The next prediction is conditioned on the preceding tokens. PaLM used a decoder-only architecture with a standard left-to-right language-modeling objective. GPT-3 was presented as an autoregressive language model.

2 · Why it exists

A sequence can be learned as a series of conditional prediction problems.

Order mattersCausal language modeling predicts the next token using only tokens on the left.
Labels are availableDuring causal language-model training, the next word serves as the label.
Generation repeatsA sequence probability can be factorized into conditional probabilities.
3 · How it works

Follow one prefix through a single prediction step.

Candidate tokens and probabilities are illustrative; the highlighted distribution is step 2.
  1. 1 · conditionThe model receives the tokens to the left of the next position.
  2. 2 · scoreThe decoder output is converted into predicted next-token probabilities.
  3. 3 · compareThe predicted distribution assigns probabilities to possible next tokens.
  4. 4 · repeatThe next prediction is conditioned on the preceding tokens.

Causal language-model training uses the observed next word as the label.

4 · Where it's used
WhoWhat they askWhat it works with
Model trainer“Which target should align with this prefix?”The observed token immediately after the prefix
Inference engineer“Which candidate tokens have the highest scores?”The next-token probability distribution
Product engineer“Why does generated text arrive one token at a time?”The repeated autoregressive decoding loop
Evaluation team“How well does the model predict held-out sequences?”The conditional likelihood of observed tokens
5 · What it solves, and what it doesn't
solves
  • Causal language-model training supplies the next word as its label.
  • A causal model can assign a probability to each next-token candidate.
  • A standard left-to-right objective predicts the next token from preceding tokens.
doesn't solve
  • A causal model cannot use tokens to the right of the position being predicted.
  • GPT-3 text samples can repeat themselves semantically.
  • GPT-3 text samples can lose coherence over sufficiently long passages.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. docsCausal language modeling, Hugging Face · read 28 Sept 2026
  2. paperImproving Language Understanding by Generative Pre-Training, OpenAI · read 28 Sept 2026
  3. paperLanguage Models are Unsupervised Multitask Learners, OpenAI · read 28 Sept 2026
  4. paperLanguage Models are Few-Shot Learners, Brown and colleagues · read 28 Sept 2026
  5. paperPaLM: Scaling Language Modeling with Pathways, Chowdhery and colleagues · read 28 Sept 2026
  6. paperAttention Is All You Need, Vaswani and colleagues · read 28 Sept 2026