Concepts

Embeddings

3 min readbeginnerUpdated 28 Sept 2026
1 · In one line

An embedding maps a discrete item such as a token ID to a dense vector of numbers.

1 · What it is

An embedding layer converts integer indexes into dense vectors of a chosen size. The Transformer used learned embeddings to convert input and output tokens to vectors. Its paper also shared the output embedding weight matrix with the pre-softmax transformation.

Word2vec computes continuous vector representations from very large data sets. Learned word vectors can place similar words near one another. BERT can combine token, segment and position information by adding their embeddings.

2 · Why it exists

Transformers convert input and output tokens to numeric vectors.

IDs are labelsAn embedding layer turns positive integer indexes into dense vectors of fixed size.
Relations matterLearned word vectors can place similar words near one another.
Inputs combineBERT forms each input representation by summing token, segment and position embeddings.
3 · How it works

Follow one token ID through an embedding lookup.

Values and positions are illustrative; the highlighted embedding conversion is step 2.
  1. 1 · identifyAn embedding layer receives positive integer indexes.
  2. 2 · look upThe embedding layer turns each index into a dense vector of fixed size.
  3. 3 · combineBERT adds token, segment and position embeddings.
  4. 4 · processThe Transformer converts input and output tokens to vectors with dimension d_model.

A word can be represented by a real-valued vector.

4 · Where it's used
WhoWhat they askWhat it works with
Language-model engineer“Which vector enters the model for this token ID?”The corresponding row in the learned embedding table
Search engineer“Which stored items have nearby vectors?”Distances between item embeddings
Model researcher“Which input components are added before BERT processes a token?”Token, segment and position embeddings
Training engineer“Which parameters are shared with the output projection?”The model's embedding and pre-softmax weight matrices
5 · What it solves, and what it doesn't
solves
  • An embedding layer converts integer indexes into dense vectors of a chosen size.
  • Learned word vectors can place similar words near one another.
  • Learned embeddings let Transformer inputs and outputs use vectors with dimension d_model.
  • BERT can combine token, segment and position information by adding their embeddings.
doesn't solve
  • An embedding lookup does not encode position or segment identity by itself; BERT adds separate position and segment embeddings.
  • A fixed embedding table cannot represent integer IDs outside its configured input range.
  • Reserving index zero for masking means zero cannot simultaneously represent a vocabulary item.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. docstf.keras.layers.Embedding, TensorFlow · read 28 Sept 2026
  2. paperEfficient Estimation of Word Representations in Vector Space, Mikolov and colleagues · read 28 Sept 2026
  3. paperGloVe: Global Vectors for Word Representation, Pennington, Socher and Manning · read 28 Sept 2026
  4. paperAttention Is All You Need, Vaswani and colleagues · read 28 Sept 2026
  5. paperBERT Pre-training of Deep Bidirectional Transformers for Language Understanding, Devlin and colleagues · read 28 Sept 2026