Concepts

Byte-pair encoding

3 min readintermediateUpdated 28 Sept 2026
1 · In one line

Byte-pair encoding builds a subword vocabulary by repeatedly merging the most frequent adjacent symbol pair.

1 · What it is

The subword adaptation of BPE starts with a character vocabulary. It counts adjacent symbol pairs, merges every occurrence of the most frequent pair into a new symbol, and repeats. Each merge expands the symbol vocabulary by one.

At test time, text is first split into the initial symbols and the learned merge operations are applied. Frequent character sequences may become single symbols, while rarer words remain sequences of smaller pieces. The merge count controls the final vocabulary size.

SentencePiece implements BPE and can train a tokenizer directly from raw sentences. GPT-2 used a byte-level BPE representation with a base vocabulary of 256 byte values. A fixed BPE model yields one subword sequence for a sentence; subword regularization instead samples multiple segmentations during training. In the character-based adaptation, a previously unseen character can still be unknown. One study found that unigram language-model tokenization matched or outperformed BPE across downstream tasks in English and Japanese.

2 · Why it exists

BPE learns variable-length symbols from adjacent-pair frequency.

Rare wordsRare words can be represented as sequences of smaller subword units.
Vocabulary sizeEach merge adds one symbol, so the number of merges controls final vocabulary size.
ReuseLearned merge operations can be applied to new text at test time.
3 · How it works

Learn two merge rules from a tiny vocabulary.

Learning byte-pair encoding merge rulesA small vocabulary is written as character symbols. Adjacent pairs are counted, the most frequent pair l plus o is merged into lo, then lo plus w is merged into low. The ordered rules are saved.
Training repeatedly counts adjacent pairs, merges the most frequent pair, and records the merge order.
  1. 1 · initializeRepresent words as sequences drawn from an initial character vocabulary.
  2. 2 · countCount adjacent symbol pairs across the training vocabulary.
  3. 3 · mergeReplace every occurrence of the most frequent pair with one new symbol.
  4. 4 · repeatContinue until the requested number of merge operations has been learned.

At test time, words are split into character sequences before the learned merge operations are applied.

4 · Where it's used
WhoWhat they askWhat it works with
Tokenizer designer“How large should the subword vocabulary be?”Initial symbols and merge count
Model engineer“How does this word become token IDs?”Ordered merge rules and vocabulary IDs
Data engineer“Was whitespace pre-tokenization used?”Tokenizer training configuration
5 · What it solves, and what it doesn't
solves
  • The adapted algorithm merges characters or character sequences rather than byte pairs.
  • It can encode open vocabularies with a fixed symbol vocabulary.
  • SentencePiece implements BPE and can train directly from raw sentences.
doesn't solve
  • The number of merge operations is still a chosen hyperparameter.
  • One fixed BPE model produces one subword sequence for a sentence.
  • Unseen characters can still be unknown in the character-based adaptation.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. paperByte Pair Encoding is Suboptimal for Language Model Pretraining, Bostrom and Durrett · read 28 Sept 2026
  2. paperNeural Machine Translation of Rare Words with Subword Units, Sennrich, Haddow and Birch · read 28 Sept 2026
  3. paperSentencePiece, Kudo and Richardson · read 28 Sept 2026
  4. paperSubword Regularization, Kudo · read 28 Sept 2026
  5. paperLanguage Models are Unsupervised Multitask Learners, OpenAI · read 28 Sept 2026