Perplexity
Perplexity is the exponentiated average negative log-likelihood that a language model assigns to a token sequence.
Perplexity starts with the probability assigned to each observed token given its preceding tokens. Convert those probabilities to negative log-likelihoods, average them, and exponentiate the mean.
One LSTM paper reported word-level perplexity on Penn Treebank and WikiText-2. The Pointer Sentinel paper reported 70.9 perplexity on Penn Treebank. GPT-2 reported exponentiated average negative log probability per canonical unit. GPT-3 calculated zero-shot perplexity on Penn Tree Bank.
Hugging Face recommends a sliding-window strategy for fixed-length models. The tokenization procedure directly affects the reported perplexity.
Perplexity uses token log-likelihoods conditioned on preceding tokens.
Average token-level negative log-likelihood, then exponentiate.
- 1 · conditionFor each position, use the probability of the observed token given preceding tokens.
- 2 · logConvert that probability to negative log-likelihood.
- 3 · averageAverage the token losses over the evaluated sequence.
- 4 · exponentiateExponentiate the mean loss to obtain perplexity.
The tokenization procedure has a direct impact on perplexity.
| Who | What they ask | What it works with |
|---|---|---|
| Language-model researcher | “How much probability did the model assign to held-out text?” | Held-out perplexity |
| Training engineer | “Did a language-model change improve evaluation loss?” | Perplexity on a fixed validation protocol |
| Benchmark reader | “Are two reported PPL values actually comparable?” | Dataset, tokenizer and windowing method |
- It summarizes autoregressive token likelihood over an evaluation sequence.
- It puts average negative log-likelihood on an exponentiated scale.
- One LSTM paper reported word-level perplexity on Penn Treebank and WikiText-2.
- It is not well defined for masked language models such as BERT.
- Disjoint fixed-length chunks can yield worse perplexity by withholding useful context.
Sources used
This explainer is written in original language. The links below support its factual claims.
- docsPerplexity of fixed-length models, Hugging Face · read 28 Sept 2026
- paperLanguage Models are Few-Shot Learners, NeurIPS / arXiv · read 28 Sept 2026
- paperPointer Sentinel Mixture Models, ICLR / arXiv · read 28 Sept 2026
- paperRegularizing and Optimizing LSTM Language Models, ICLR / arXiv · read 28 Sept 2026
- paperLanguage Models are Unsupervised Multitask Learners, OpenAI · read 28 Sept 2026