Tokens
Tokens are chunks of text processed by text-generation and embedding models.
A text token is a character sequence that a tokenizer maps through a vocabulary. It may cover one character, punctuation, a word or part of a word. The printed piece and the token ID are related by one particular encoding, not by a universal numbering system.
Before a Transformer processes text, learned embeddings convert its input tokens into vectors. Position information is then added because the architecture otherwise has no recurrence or convolution to express order. At the output, a learned transformation and softmax produce probabilities for the next token. The selected token can then join the sequence and become input for the following prediction.
The same visible text can produce different token counts depending on the model, encoding and language. OpenAI also distinguishes input, output, cached input and internal reasoning tokens for accounting.
Neural text-generation systems can use a vocabulary whose size is fixed before model training.
Follow one token from printed text to a next-token choice.
- 1 · segmentThe tokenizer divides the string into pieces from its vocabulary.
- 2 · identifyVocabulary lookup assigns each piece its integer token ID.
- 3 · embedThe embedding table converts each token ID to a learned vector.
- 4 · predictThe model converts its output vector into probabilities over possible next tokens.
A token is not inherently a word: the chosen encoding decides whether a piece is a character, punctuation mark, word or subword.
| Who | What they ask | What it works with |
|---|---|---|
| API developer | “How many input and output tokens will this request use?” | The target model's encoding and token counters |
| Model engineer | “Which embedding row represents this text piece?” | The tokenizer vocabulary and token ID |
| Product designer | “How much conversation fits in the context window?” | Input, cached, reasoning and output token usage |
| Localization team | “Why is this translation longer in tokens?” | The model's encoding for each language |
- A token vocabulary lets a tokenizer convert text into a sequence of IDs.
- Subword tokens let a model represent rare words using smaller known pieces.
- Token IDs select learned representations that embedding layers convert to vectors.
- Cached input tokens reuse earlier input, and their pricing can differ from uncached input.
- A token is not a universal unit of meaning; boundaries depend on the encoding.
- A token count is not a word count, and the ratio varies across languages and encodings.
- An example token ID from one encoding must not be assumed to apply to another model.
- Tokens are discrete inputs; learned embeddings still convert them to vectors for model processing.
Sources used
This explainer is written in original language. The links below support its factual claims.
- officialUnderstanding and counting tokens, OpenAI · read 27 Sept 2026
- docsKey concepts, OpenAI · read 27 Sept 2026
- paperAttention Is All You Need, Vaswani et al. · read 27 Sept 2026
- paperSentencePiece: A simple and language independent subword tokenizer and detokenizer for Neural Text Processing, Kudo and Richardson · read 27 Sept 2026
- paperNeural Machine Translation of Rare Words with Subword Units, Sennrich, Haddow and Birch · read 27 Sept 2026