Concepts

Attention

3 min readbeginnerUpdated 28 Sept 2026
1 · In one line

Attention lets a model build an output by assigning weights to available pieces of information and mixing their values.

1 · What it is

Attention is a learned way to retrieve from a set of representations. Attention maps a query and a set of key-value pairs to an output.

The model compares the query with the keys. Softmax converts the compatibility scores into weights on the values. The output is computed as a weighted sum of the values.

Source-sentence representations can be selectively retrieved by the decoder. Attention has also been used to select image regions while a model generates caption words. Scaled dot-product attention divides query-key dot products by √dk and applies softmax. PyTorch exposes scaled dot-product attention with query, key, value and an optional attention mask as inputs.

2 · Why it exists

One fixed summary can lose information that a later prediction needs.

Fixed bottleneckEarly encoder-decoder translation models compressed a whole source sentence into one fixed-length vector.
Changing needThe useful source words can differ for each word the decoder is about to produce.
Hard selectionSoft attention keeps the selection differentiable, so the alignment model can train with the rest of the network.
3 · How it works

Follow one query as it reads three stored values.

Attention normalizes query-key scores into weights on the values.
  1. 1 · queryAttention maps a query and a set of key-value pairs to an output.
  2. 2 · scoreA compatibility function compares that query with every available key.
  3. 3 · weightSoftmax turns the scores into weights used to mix the values.
  4. 4 · mixThe output is the weighted sum of the values paired with those keys.

The weights are recomputed for each query, so the same memory can yield a different context next time.

4 · Where it's used
WhoWhat they askWhat it works with
Translation model“Which source words matter for the next translated word?”Encoded positions in the source sentence
Image captioner“Which image regions matter for the next word?”Feature vectors for image regions
Transformer layer“Which token representations should update this position?”Keys and values supplied to the layer
5 · What it solves, and what it doesn't
solves
  • Attention removes the need to squeeze every source position into one fixed context vector.
  • It produces a context vector as a weighted sum of source representations.
  • Global attention can consider every source state, while local attention can limit the comparison to a smaller window.
doesn't solve
  • Global attention still compares against every available source position for each target position.
  • Local attention examines only a subset of source positions rather than every source word.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. paperNeural Machine Translation by Jointly Learning to Align and Translate, Bahdanau, Cho and Bengio · read 27 Sept 2026
  2. paperEffective Approaches to Attention-based Neural Machine Translation, Luong, Pham and Manning · read 27 Sept 2026
  3. paperAttention Is All You Need, Vaswani et al. · read 27 Sept 2026
  4. docstorch.nn.functional.scaled_dot_product_attention, PyTorch · read 27 Sept 2026
  5. paperShow, Attend and Tell Neural Image Caption Generation with Visual Attention, Xu et al. · read 27 Sept 2026