Perplexity (metric)Concepts

Perplexity

4 min readintermediateUpdated 28 Sept 2026
1 · In one line

Perplexity is the exponentiated average negative log-likelihood that a language model assigns to a token sequence.

1 · What it is

Perplexity starts with the probability assigned to each observed token given its preceding tokens. Convert those probabilities to negative log-likelihoods, average them, and exponentiate the mean.

One LSTM paper reported word-level perplexity on Penn Treebank and WikiText-2. The Pointer Sentinel paper reported 70.9 perplexity on Penn Treebank. GPT-2 reported exponentiated average negative log probability per canonical unit. GPT-3 calculated zero-shot perplexity on Penn Tree Bank.

Hugging Face recommends a sliding-window strategy for fixed-length models. The tokenization procedure directly affects the reported perplexity.

2 · Why it exists

Perplexity uses token log-likelihoods conditioned on preceding tokens.

Sequence scoreIt summarizes conditional token probabilities across a sequence.
Evaluation conventionFixed-length models need a declared windowing strategy.
Tokenizer dependenceTokenization directly affects the reported value.
3 · How it works

Average token-level negative log-likelihood, then exponentiate.

Toy arithmetic for PPL = exp(mean negative log-likelihood). Tokenization and fixed-window evaluation procedure affect the reported value.
  1. 1 · conditionFor each position, use the probability of the observed token given preceding tokens.
  2. 2 · logConvert that probability to negative log-likelihood.
  3. 3 · averageAverage the token losses over the evaluated sequence.
  4. 4 · exponentiateExponentiate the mean loss to obtain perplexity.

The tokenization procedure has a direct impact on perplexity.

4 · Where it's used
WhoWhat they askWhat it works with
Language-model researcher“How much probability did the model assign to held-out text?”Held-out perplexity
Training engineer“Did a language-model change improve evaluation loss?”Perplexity on a fixed validation protocol
Benchmark reader“Are two reported PPL values actually comparable?”Dataset, tokenizer and windowing method
5 · What it solves, and what it doesn't
solves
  • It summarizes autoregressive token likelihood over an evaluation sequence.
  • It puts average negative log-likelihood on an exponentiated scale.
  • One LSTM paper reported word-level perplexity on Penn Treebank and WikiText-2.
doesn't solve
  • It is not well defined for masked language models such as BERT.
  • Disjoint fixed-length chunks can yield worse perplexity by withholding useful context.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. docsPerplexity of fixed-length models, Hugging Face · read 28 Sept 2026
  2. paperLanguage Models are Few-Shot Learners, NeurIPS / arXiv · read 28 Sept 2026
  3. paperPointer Sentinel Mixture Models, ICLR / arXiv · read 28 Sept 2026
  4. paperRegularizing and Optimizing LSTM Language Models, ICLR / arXiv · read 28 Sept 2026
  5. paperLanguage Models are Unsupervised Multitask Learners, OpenAI · read 28 Sept 2026