Concepts

Multi-head attention

3 min readintermediateUpdated 28 Sept 2026
1 · In one line

Multi-head attention consists of several attention layers running in parallel.

1 · What it is

Multi-head attention duplicates the attention calculation into several heads. Each head receives different learned projections of the queries, keys and values.

The projected queries, keys and values pass through attention in parallel. Its output is one weighted mixture of its projected values for each query. The head outputs are concatenated and projected to produce the final values.

The heads run in parallel. The heads use different learned linear projections. PyTorch’s MultiheadAttention is a reference implementation of the original architecture. TensorFlow’s implementation concatenates values after interpolating them with attention probabilities. BERT’s model architecture is a multi-layer bidirectional Transformer encoder.

2 · Why it exists

One attention calculation compresses its matches into one weighted average.

One projectionA single head applies one attention function to its queries, keys and values.
Averaged evidenceOne head merges its selected values into one weighted sum for each query.
Different relationsMultiple heads give the layer several learned projections for attending at different positions and representation subspaces.
3 · How it works

Follow one sequence through three attention heads.

Each head has its own Q, K and V projections; concatenation and an output projection return one combined representation.
  1. 1 · projectEvery head uses its own learned linear projections of the queries, keys and values.
  2. 2 · attendThe attention functions for the heads run in parallel.
  3. 3 · joinThe per-head outputs are concatenated into one wider vector.
  4. 4 · projectA final learned output projection mixes the joined head outputs back into the model dimension.

Each head uses different learned projections of the queries, keys and values.

4 · Where it's used
WhoWhat they askWhat it works with
Transformer encoder“Which positions should update each input representation?”Representations from the same encoder layer
Transformer decoder“Which earlier positions matter for this position?”Causally permitted decoder states
Encoder-decoder model“Which encoded input positions matter while decoding?”Encoder outputs used as keys and values
5 · What it solves, and what it doesn't
solves
  • Separate learned projections let the layer attend from several representation subspaces.
  • The attention functions for the heads run in parallel.
  • Concatenation preserves every head's output until the learned output projection combines them.
doesn't solve
  • Multi-head attention still inherits the pairwise score matrix of full attention.
  • The mechanism does not add sequence order unless position information is supplied elsewhere.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. paperAttention Is All You Need, Vaswani et al. · read 27 Sept 2026
  2. docsMultiheadAttention, PyTorch · read 27 Sept 2026
  3. docstorch.nn.functional.scaled_dot_product_attention, PyTorch · read 27 Sept 2026
  4. paperBERT Pre-training of Deep Bidirectional Transformers for Language Understanding, Devlin et al. · read 27 Sept 2026
  5. docstf.keras.layers.MultiHeadAttention, TensorFlow · read 27 Sept 2026