VocabularyConcepts
Model vocabulary
1 · In one line
A model vocabulary supplies the mapping used to convert text into an ID sequence and back.
1 · What it is
SentencePiece manages a vocabulary-to-ID mapping that converts text into an ID sequence and back. A tokenizer maps selected pieces to their corresponding IDs in the model vocabulary.
A fixed vocabulary does not make every token a complete word. Subword units can encode rare and unknown words. BERT used a WordPiece vocabulary containing 30,000 tokens. GPT-2 expanded its byte-level BPE vocabulary to 50,257 entries.
SentencePiece sets the vocabulary size before neural-model training.
Text keeps growingSubword units can encode rare and unknown words.
Models use IDsA tokenizer maps its selected pieces to their corresponding vocabulary IDs.
Size is chosenSentencePiece sets the vocabulary size before the neural model is trained.
Follow one word from characters to vocabulary IDs.
- 1 · splitA subword tokenizer can represent an unfamiliar word as a sequence of known pieces.
- 2 · look upThe vocabulary maps each selected piece to its corresponding integer ID.
- 3 · passThe resulting ID sequence becomes the tokenizer output supplied to the model.
- 4 · decodeThe same mapping can convert an ID sequence back into text pieces.
A vocabulary fixes the available pieces; it does not require every piece to be a whole word.
| Who | What they ask | What it works with |
|---|---|---|
| Tokenizer designer | “How many pieces should this vocabulary contain?” | Candidate pieces and the chosen vocabulary size |
| Model trainer | “Which integer represents this text piece?” | The vocabulary-to-ID mapping stored with the tokenizer |
| Application developer | “Why did this word become several IDs?” | The model's tokenizer vocabulary and segmentation |
| Debugging team | “Which special IDs mark sequence boundaries?” | The tokenizer's reserved meta symbols |
solves
- A fixed vocabulary gives a tokenizer a predetermined number of pieces.
- Subword pieces can represent rare or unknown words as sequences instead of one unknown word.
- The vocabulary-to-ID mapping converts selected pieces into an integer sequence and back.
- Reserved vocabulary IDs can represent unknown, beginning, end and padding symbols.
doesn't solve
- A fixed vocabulary does not make every token a complete word.
- A vocabulary size does not specify how a tokenizer chooses pieces within a string.
- Vocabulary IDs correspond to the vocabulary of a particular model.
6 · Go deeper
Sources used
This explainer is written in original language. The links below support its factual claims.
- paperSentencePiece: A simple and language independent subword tokenizer and detokenizer for Neural Text Processing, Kudo and Richardson · read 28 Sept 2026
- paperNeural Machine Translation of Rare Words with Subword Units, Sennrich, Haddow and Birch · read 28 Sept 2026
- paperBERT Pre-training of Deep Bidirectional Transformers for Language Understanding, Devlin and colleagues · read 28 Sept 2026
- docsThe tokenization pipeline, Hugging Face · read 28 Sept 2026
- paperLanguage Models are Unsupervised Multitask Learners, OpenAI · read 28 Sept 2026