Concepts

Embedding layer

3 min readbeginnerUpdated 28 Sept 2026
1 · In one line

An embedding layer turns integer indexes into dense vectors of fixed size.

1 · What it is

An embedding layer turns integer indexes into dense vectors of fixed size. PyTorch implements it as a lookup table. Its weight matrix is learnable.

The input tensor contains indexes to extract from the embedding weight matrix. The output keeps the input shape and appends the embedding dimension. Entries at a specified padding index do not contribute to the gradient.

Word2vec learns word vectors from large data sets. GloVe uses a global log-bilinear regression model for unsupervised word representations. Transformers inject relative or absolute token-position information.

2 · Why it exists

Embedding layers turn discrete indexes into vectors.

Integer inputTensorFlow's embedding layer accepts integer indexes.
Dense outputIt turns those indexes into dense vectors of fixed size.
Learned rowsPyTorch describes the embedding matrix as a learnable weight.
3 · How it works

Follow three token IDs through one embedding lookup.

The input indexes identify entries to extract from the embedding weight matrix.
  1. 1 · indexSupply a tensor of integer indexes.
  2. 2 · lookupUse the supplied indexes to extract entries from the embedding weight matrix.
  3. 3 · orderThe output retains the input shape and appends the embedding dimension.

The supplied indexes identify entries to extract from the embedding weight matrix.

4 · Where it's used
WhoWhat they askWhat it works with
Language-model team“Which vector represents each token ID?”Rows in the token embedding matrix
Recommender team“Which vector represents this item ID?”Rows indexed by item identifiers
Feature engineer“How should categorical IDs enter a neural network?”Learned category vectors
5 · What it solves, and what it doesn't
solves
  • PyTorch implements embeddings as a lookup table with a fixed dictionary size and embedding size.
  • TensorFlow returns an output with one extra embedding dimension.
  • Word2vec learns word vectors from large data sets.
doesn't solve
  • The input indexes must stay within the configured vocabulary range.
  • Entries at a specified padding index do not contribute to the gradient.
  • Transformers inject relative or absolute token-position information.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. docsEmbedding, PyTorch · read 28 Sept 2026
  2. docstf.keras.layers.Embedding, TensorFlow · read 28 Sept 2026
  3. paperEfficient Estimation of Word Representations in Vector Space, Mikolov et al. · read 28 Sept 2026
  4. paperGloVe, Pennington, Socher and Manning · read 28 Sept 2026
  5. paperAttention Is All You Need, Vaswani et al. · read 28 Sept 2026