Decoder-only modelsConcepts
Decoder-only transformer models
1 · In one line
A decoder-only model predicts each next token from the tokens that came before it.
1 · What it is
The first GPT used a multi-layer Transformer decoder for language modeling. It used masked self-attention heads. That mask keeps a position from reading tokens to its right.
GPT-2 was trained to predict the next word from previous words.
GPT-3 is an autoregressive language model. PaLM is a densely activated Transformer language model.
A causal language model can attend only to tokens on the left.
Future leakageA causal language model can attend only to tokens on the left.
Repeated choiceGPT-2 predicts the next word from all previous words.
Shared interfaceA prompt and demonstrations can specify a task through text interaction.
Follow a prompt through one next-token step.
- 1 · prefixThe model receives the tokens generated or supplied so far.
- 2 · maskMasked self-attention blocks information from subsequent positions.
- 3 · scoreThe final decoder state produces a distribution over target tokens.
- 4 · nextCausal language modeling predicts the next token in a sequence.
GPT-3 is an autoregressive language model.
| Who | What they ask | What it works with |
|---|---|---|
| Writing assistant | “What text could continue this paragraph?” | The prompt and generated prefix |
| Code assistant | “Which token should follow this source code?” | Code tokens already present |
| Few-shot system | “What output follows these examples and this new input?” | Instructions, examples and the unfinished answer |
solves
- A causal language model can attend only to tokens on the left.
- GPT-2 can generate a lengthy continuation from an input.
- GPT-3 can receive tasks and demonstrations through text interaction.
doesn't solve
- GPT-2 exhibits world modeling failures.
- GPT-2 can produce repetitive text and unnatural topic changes.
6 · Go deeper
Sources used
This explainer is written in original language. The links below support its factual claims.
- paperImproving Language Understanding by Generative Pre-Training, Radford et al., OpenAI · read 28 Sept 2026
- officialBetter language models and their implications, OpenAI · read 28 Sept 2026
- paperLanguage Models are Few-Shot Learners, Brown et al. · read 28 Sept 2026
- paperPaLM Scaling Language Modeling with Pathways, Chowdhery et al. · read 28 Sept 2026
- docsCausal language modeling, Hugging Face · read 28 Sept 2026