Embedding layer
An embedding layer turns integer indexes into dense vectors of fixed size.
An embedding layer turns integer indexes into dense vectors of fixed size. PyTorch implements it as a lookup table. Its weight matrix is learnable.
The input tensor contains indexes to extract from the embedding weight matrix. The output keeps the input shape and appends the embedding dimension. Entries at a specified padding index do not contribute to the gradient.
Word2vec learns word vectors from large data sets. GloVe uses a global log-bilinear regression model for unsupervised word representations. Transformers inject relative or absolute token-position information.
Embedding layers turn discrete indexes into vectors.
Follow three token IDs through one embedding lookup.
- 1 · indexSupply a tensor of integer indexes.
- 2 · lookupUse the supplied indexes to extract entries from the embedding weight matrix.
- 3 · orderThe output retains the input shape and appends the embedding dimension.
The supplied indexes identify entries to extract from the embedding weight matrix.
| Who | What they ask | What it works with |
|---|---|---|
| Language-model team | “Which vector represents each token ID?” | Rows in the token embedding matrix |
| Recommender team | “Which vector represents this item ID?” | Rows indexed by item identifiers |
| Feature engineer | “How should categorical IDs enter a neural network?” | Learned category vectors |
- PyTorch implements embeddings as a lookup table with a fixed dictionary size and embedding size.
- TensorFlow returns an output with one extra embedding dimension.
- Word2vec learns word vectors from large data sets.
- The input indexes must stay within the configured vocabulary range.
- Entries at a specified padding index do not contribute to the gradient.
- Transformers inject relative or absolute token-position information.
Sources used
This explainer is written in original language. The links below support its factual claims.
- docsEmbedding, PyTorch · read 28 Sept 2026
- docstf.keras.layers.Embedding, TensorFlow · read 28 Sept 2026
- paperEfficient Estimation of Word Representations in Vector Space, Mikolov et al. · read 28 Sept 2026
- paperGloVe, Pennington, Socher and Manning · read 28 Sept 2026
- paperAttention Is All You Need, Vaswani et al. · read 28 Sept 2026