Byte-pair encoding
Byte-pair encoding builds a subword vocabulary by repeatedly merging the most frequent adjacent symbol pair.
The subword adaptation of BPE starts with a character vocabulary. It counts adjacent symbol pairs, merges every occurrence of the most frequent pair into a new symbol, and repeats. Each merge expands the symbol vocabulary by one.
At test time, text is first split into the initial symbols and the learned merge operations are applied. Frequent character sequences may become single symbols, while rarer words remain sequences of smaller pieces. The merge count controls the final vocabulary size.
SentencePiece implements BPE and can train a tokenizer directly from raw sentences. GPT-2 used a byte-level BPE representation with a base vocabulary of 256 byte values. A fixed BPE model yields one subword sequence for a sentence; subword regularization instead samples multiple segmentations during training. In the character-based adaptation, a previously unseen character can still be unknown. One study found that unigram language-model tokenization matched or outperformed BPE across downstream tasks in English and Japanese.
BPE learns variable-length symbols from adjacent-pair frequency.
Learn two merge rules from a tiny vocabulary.
- 1 · initializeRepresent words as sequences drawn from an initial character vocabulary.
- 2 · countCount adjacent symbol pairs across the training vocabulary.
- 3 · mergeReplace every occurrence of the most frequent pair with one new symbol.
- 4 · repeatContinue until the requested number of merge operations has been learned.
At test time, words are split into character sequences before the learned merge operations are applied.
| Who | What they ask | What it works with |
|---|---|---|
| Tokenizer designer | “How large should the subword vocabulary be?” | Initial symbols and merge count |
| Model engineer | “How does this word become token IDs?” | Ordered merge rules and vocabulary IDs |
| Data engineer | “Was whitespace pre-tokenization used?” | Tokenizer training configuration |
- The adapted algorithm merges characters or character sequences rather than byte pairs.
- It can encode open vocabularies with a fixed symbol vocabulary.
- SentencePiece implements BPE and can train directly from raw sentences.
- The number of merge operations is still a chosen hyperparameter.
- One fixed BPE model produces one subword sequence for a sentence.
- Unseen characters can still be unknown in the character-based adaptation.
Sources used
This explainer is written in original language. The links below support its factual claims.
- paperByte Pair Encoding is Suboptimal for Language Model Pretraining, Bostrom and Durrett · read 28 Sept 2026
- paperNeural Machine Translation of Rare Words with Subword Units, Sennrich, Haddow and Birch · read 28 Sept 2026
- paperSentencePiece, Kudo and Richardson · read 28 Sept 2026
- paperSubword Regularization, Kudo · read 28 Sept 2026
- paperLanguage Models are Unsupervised Multitask Learners, OpenAI · read 28 Sept 2026