Concepts

Tokens

3 min readbeginnerUpdated 28 Sept 2026
1 · In one line

Tokens are chunks of text processed by text-generation and embedding models.

1 · What it is

A text token is a character sequence that a tokenizer maps through a vocabulary. It may cover one character, punctuation, a word or part of a word. The printed piece and the token ID are related by one particular encoding, not by a universal numbering system.

Before a Transformer processes text, learned embeddings convert its input tokens into vectors. Position information is then added because the architecture otherwise has no recurrence or convolution to express order. At the output, a learned transformation and softmax produce probabilities for the next token. The selected token can then join the sequence and become input for the following prediction.

The same visible text can produce different token counts depending on the model, encoding and language. OpenAI also distinguishes input, output, cached input and internal reasoning tokens for accounting.

2 · Why it exists

Neural text-generation systems can use a vocabulary whose size is fixed before model training.

Text is variablePrinted text can contain characters, punctuation, complete words and word fragments.
Networks use vectorsAn embedding layer converts each input token into a learned vector before the model processes the sequence.
Output is sequentialA generative language model produces output tokens, one position after another.
3 · How it works

Follow one token from printed text to a next-token choice.

The example pieces and symbolic IDs are illustrative; the highlighted box is step 3.
  1. 1 · segmentThe tokenizer divides the string into pieces from its vocabulary.
  2. 2 · identifyVocabulary lookup assigns each piece its integer token ID.
  3. 3 · embedThe embedding table converts each token ID to a learned vector.
  4. 4 · predictThe model converts its output vector into probabilities over possible next tokens.

A token is not inherently a word: the chosen encoding decides whether a piece is a character, punctuation mark, word or subword.

4 · Where it's used
WhoWhat they askWhat it works with
API developer“How many input and output tokens will this request use?”The target model's encoding and token counters
Model engineer“Which embedding row represents this text piece?”The tokenizer vocabulary and token ID
Product designer“How much conversation fits in the context window?”Input, cached, reasoning and output token usage
Localization team“Why is this translation longer in tokens?”The model's encoding for each language
5 · What it solves, and what it doesn't
solves
  • A token vocabulary lets a tokenizer convert text into a sequence of IDs.
  • Subword tokens let a model represent rare words using smaller known pieces.
  • Token IDs select learned representations that embedding layers convert to vectors.
  • Cached input tokens reuse earlier input, and their pricing can differ from uncached input.
doesn't solve
  • A token is not a universal unit of meaning; boundaries depend on the encoding.
  • A token count is not a word count, and the ratio varies across languages and encodings.
  • An example token ID from one encoding must not be assumed to apply to another model.
  • Tokens are discrete inputs; learned embeddings still convert them to vectors for model processing.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. officialUnderstanding and counting tokens, OpenAI · read 27 Sept 2026
  2. docsKey concepts, OpenAI · read 27 Sept 2026
  3. paperAttention Is All You Need, Vaswani et al. · read 27 Sept 2026
  4. paperSentencePiece: A simple and language independent subword tokenizer and detokenizer for Neural Text Processing, Kudo and Richardson · read 27 Sept 2026
  5. paperNeural Machine Translation of Rare Words with Subword Units, Sennrich, Haddow and Birch · read 27 Sept 2026