Tokenization
Tokenization turns text into a sequence of vocabulary IDs that a model can process, then reverses generated IDs back into readable text.
A tokenization pipeline may normalize text, pre-tokenize it, apply a segmentation model and add special tokens. Normalization can change the raw string before the tokenizer chooses pieces. Pre-tokenization can impose boundaries such as whitespace and punctuation before the trained model selects final pieces. The segmentation model then applies learned rules and maps each selected piece to a vocabulary ID. Post-processing may add special tokens required by a particular model.
Subword tokenization keeps common pieces while allowing rare or unseen words to be expressed as several pieces. SentencePiece can train a subword model directly from raw sentences rather than requiring word-level pre-tokenization. WordPiece instead pre-tokenizes on whitespace and punctuation, then greedily chooses vocabulary pieces within each word. Different valid tokenizers can therefore produce different ID sequences for the same printed string.
OpenAI notes that spaces and capitalization can change how a word is divided. The same text can also have different token counts across models, encodings and languages. That difference matters because a model’s prompt and generated output share a maximum context length.
Neural text-generation systems can use a vocabulary whose size is fixed before model training.
Follow one short string from characters to model input.
- 1 · normalizeOptional normalization applies the tokenizer's recorded text rules before segmentation.
- 2 · splitA trained tokenization model chooses pieces and maps them to IDs in its vocabulary.
- 3 · mapThe vocabulary maps every chosen piece to an integer ID.
- 4 · decodeFor generated output, the decoder maps IDs back to pieces and joins them into text.
Changing a tokenizer's normalization can require retraining it.
| Who | What they ask | What it works with |
|---|---|---|
| API developer | “Will this prompt fit the model's context limit?” | The tokenizer and model-specific token count |
| Model trainer | “Which text pieces should the vocabulary contain?” | Frequencies and candidate pieces in the training corpus |
| Search engineer | “How will this document become embedding-model input?” | The embedding model's tokenizer and maximum input length |
| Localization team | “Why does this language use more tokens for the same message?” | The model's encoding and the language's segmentation |
- A fixed vocabulary gives the model a finite set of input and output symbols.
- Subword methods can represent an unfamiliar word as several known pieces instead of one unknown item.
- A self-contained SentencePiece model stores its vocabulary, segmentation parameters and a compiled character normalizer.
- Decoding can turn generated token IDs back into readable text.
- Token boundaries do not guarantee linguistic boundaries; a token may be a character, punctuation mark, word or part of a word.
- Token counts are not portable between model encodings, even when the visible text is identical.
- Tokenization prepares a sequence of discrete elements for a machine-learning model; later model layers process that sequence.
- An example token ID from one encoding must not be assumed to apply to another model.
Sources used
This explainer is written in original language. The links below support its factual claims.
- officialUnderstanding and counting tokens, OpenAI · read 27 Sept 2026
- docsThe tokenization pipeline, Hugging Face · read 27 Sept 2026
- paperSentencePiece: A simple and language independent subword tokenizer and detokenizer for Neural Text Processing, Kudo and Richardson · read 27 Sept 2026
- paperNeural Machine Translation of Rare Words with Subword Units, Sennrich, Haddow and Birch · read 27 Sept 2026
- officialA Fast WordPiece Tokenization System, Google Research · read 27 Sept 2026