Concepts
Autoregressive models
1 · In one line
An autoregressive model factorizes a sequence probability into conditional probabilities.
1 · What it is
GPT-2 expresses a sequence probability as a product of conditional probabilities.
PixelRNN predicts image pixels sequentially along two spatial dimensions. WaveNet is fully probabilistic and autoregressive.
XLNet maximizes expected likelihood over permutations of the factorization order.
A joint sequence can be turned into a chain of smaller prediction problems.
Ordered dependencePixelRNN predicts image pixels sequentially across two spatial dimensions.
Exact likelihoodAutoregressive language models make likelihoods and sampling straightforward to compute.
Many mediaWaveNet conditions each audio sample on the samples before it.
Follow one sequence through three conditional predictions.
- 1 · orderChoose an order for the elements in the sequence.
- 2 · conditionModel the next element from the elements that precede it.
- 3 · factorWrite one conditional probability for each element in the chosen order.
- 4 · multiplyRepresent the joint sequence probability as the product of those conditional probabilities.
XLNet considers permutations of the factorization order.
| Who | What they ask | What it works with |
|---|---|---|
| Language-model researcher | “What distribution should the decoder learn at each text position?” | Conditional probabilities over the next token |
| Audio researcher | “What should determine the next waveform sample?” | The audio samples already generated |
| Image researcher | “How can an image be treated as an ordered prediction task?” | Earlier pixels in the chosen scan order |
| Evaluation engineer | “What probability does the model assign to this observed sequence?” | The product of its conditional probabilities |
solves
- A joint sequence probability can be factorized into conditional probabilities.
- PixelRNN applies sequential prediction to image pixels.
- Autoregressive language models support direct likelihood calculation and sampling.
doesn't solve
- GPT-3 samples can lose coherence over sufficiently long passages.
6 · Go deeper
Sources used
This explainer is written in original language. The links below support its factual claims.
- paperLanguage Models are Unsupervised Multitask Learners, OpenAI · read 28 Sept 2026
- paperPixel Recurrent Neural Networks, van den Oord, Kalchbrenner and Kavukcuoglu · read 28 Sept 2026
- paperWaveNet: A Generative Model for Raw Audio, van den Oord and colleagues · read 28 Sept 2026
- paperLanguage Models are Few-Shot Learners, Brown and colleagues · read 28 Sept 2026
- paperXLNet: Generalized Autoregressive Pretraining for Language Understanding, Yang and colleagues · read 28 Sept 2026