Concepts
Embeddings
1 · In one line
An embedding maps a discrete item such as a token ID to a dense vector of numbers.
1 · What it is
An embedding layer converts integer indexes into dense vectors of a chosen size. The Transformer used learned embeddings to convert input and output tokens to vectors. Its paper also shared the output embedding weight matrix with the pre-softmax transformation.
Word2vec computes continuous vector representations from very large data sets. Learned word vectors can place similar words near one another. BERT can combine token, segment and position information by adding their embeddings.
Transformers convert input and output tokens to numeric vectors.
IDs are labelsAn embedding layer turns positive integer indexes into dense vectors of fixed size.
Relations matterLearned word vectors can place similar words near one another.
Inputs combineBERT forms each input representation by summing token, segment and position embeddings.
Follow one token ID through an embedding lookup.
- 1 · identifyAn embedding layer receives positive integer indexes.
- 2 · look upThe embedding layer turns each index into a dense vector of fixed size.
- 3 · combineBERT adds token, segment and position embeddings.
- 4 · processThe Transformer converts input and output tokens to vectors with dimension d_model.
A word can be represented by a real-valued vector.
| Who | What they ask | What it works with |
|---|---|---|
| Language-model engineer | “Which vector enters the model for this token ID?” | The corresponding row in the learned embedding table |
| Search engineer | “Which stored items have nearby vectors?” | Distances between item embeddings |
| Model researcher | “Which input components are added before BERT processes a token?” | Token, segment and position embeddings |
| Training engineer | “Which parameters are shared with the output projection?” | The model's embedding and pre-softmax weight matrices |
solves
- An embedding layer converts integer indexes into dense vectors of a chosen size.
- Learned word vectors can place similar words near one another.
- Learned embeddings let Transformer inputs and outputs use vectors with dimension d_model.
- BERT can combine token, segment and position information by adding their embeddings.
doesn't solve
- An embedding lookup does not encode position or segment identity by itself; BERT adds separate position and segment embeddings.
- A fixed embedding table cannot represent integer IDs outside its configured input range.
- Reserving index zero for masking means zero cannot simultaneously represent a vocabulary item.
6 · Go deeper
Sources used
This explainer is written in original language. The links below support its factual claims.
- docstf.keras.layers.Embedding, TensorFlow · read 28 Sept 2026
- paperEfficient Estimation of Word Representations in Vector Space, Mikolov and colleagues · read 28 Sept 2026
- paperGloVe: Global Vectors for Word Representation, Pennington, Socher and Manning · read 28 Sept 2026
- paperAttention Is All You Need, Vaswani and colleagues · read 28 Sept 2026
- paperBERT Pre-training of Deep Bidirectional Transformers for Language Understanding, Devlin and colleagues · read 28 Sept 2026