Decoder-only modelsConcepts

Decoder-only transformer models

3 min readbeginnerUpdated 28 Sept 2026
1 · In one line

A decoder-only model predicts each next token from the tokens that came before it.

1 · What it is

The first GPT used a multi-layer Transformer decoder for language modeling. It used masked self-attention heads. That mask keeps a position from reading tokens to its right.

GPT-2 was trained to predict the next word from previous words.

GPT-3 is an autoregressive language model. PaLM is a densely activated Transformer language model.

2 · Why it exists

A causal language model can attend only to tokens on the left.

Future leakageA causal language model can attend only to tokens on the left.
Repeated choiceGPT-2 predicts the next word from all previous words.
Shared interfaceA prompt and demonstrations can specify a task through text interaction.
3 · How it works

Follow a prompt through one next-token step.

The model can attend only to tokens on the left. It predicts the next token in the sequence.
  1. 1 · prefixThe model receives the tokens generated or supplied so far.
  2. 2 · maskMasked self-attention blocks information from subsequent positions.
  3. 3 · scoreThe final decoder state produces a distribution over target tokens.
  4. 4 · nextCausal language modeling predicts the next token in a sequence.

GPT-3 is an autoregressive language model.

4 · Where it's used
WhoWhat they askWhat it works with
Writing assistant“What text could continue this paragraph?”The prompt and generated prefix
Code assistant“Which token should follow this source code?”Code tokens already present
Few-shot system“What output follows these examples and this new input?”Instructions, examples and the unfinished answer
5 · What it solves, and what it doesn't
solves
  • A causal language model can attend only to tokens on the left.
  • GPT-2 can generate a lengthy continuation from an input.
  • GPT-3 can receive tasks and demonstrations through text interaction.
doesn't solve
  • GPT-2 exhibits world modeling failures.
  • GPT-2 can produce repetitive text and unnatural topic changes.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. paperImproving Language Understanding by Generative Pre-Training, Radford et al., OpenAI · read 28 Sept 2026
  2. officialBetter language models and their implications, OpenAI · read 28 Sept 2026
  3. paperLanguage Models are Few-Shot Learners, Brown et al. · read 28 Sept 2026
  4. paperPaLM Scaling Language Modeling with Pathways, Chowdhery et al. · read 28 Sept 2026
  5. docsCausal language modeling, Hugging Face · read 28 Sept 2026