Attention
Attention lets a model build an output by assigning weights to available pieces of information and mixing their values.
Attention is a learned way to retrieve from a set of representations. Attention maps a query and a set of key-value pairs to an output.
The model compares the query with the keys. Softmax converts the compatibility scores into weights on the values. The output is computed as a weighted sum of the values.
Source-sentence representations can be selectively retrieved by the decoder. Attention has also been used to select image regions while a model generates caption words. Scaled dot-product attention divides query-key dot products by √dk and applies softmax. PyTorch exposes scaled dot-product attention with query, key, value and an optional attention mask as inputs.
One fixed summary can lose information that a later prediction needs.
Follow one query as it reads three stored values.
- 1 · queryAttention maps a query and a set of key-value pairs to an output.
- 2 · scoreA compatibility function compares that query with every available key.
- 3 · weightSoftmax turns the scores into weights used to mix the values.
- 4 · mixThe output is the weighted sum of the values paired with those keys.
The weights are recomputed for each query, so the same memory can yield a different context next time.
| Who | What they ask | What it works with |
|---|---|---|
| Translation model | “Which source words matter for the next translated word?” | Encoded positions in the source sentence |
| Image captioner | “Which image regions matter for the next word?” | Feature vectors for image regions |
| Transformer layer | “Which token representations should update this position?” | Keys and values supplied to the layer |
- Attention removes the need to squeeze every source position into one fixed context vector.
- It produces a context vector as a weighted sum of source representations.
- Global attention can consider every source state, while local attention can limit the comparison to a smaller window.
- Global attention still compares against every available source position for each target position.
- Local attention examines only a subset of source positions rather than every source word.
Sources used
This explainer is written in original language. The links below support its factual claims.
- paperNeural Machine Translation by Jointly Learning to Align and Translate, Bahdanau, Cho and Bengio · read 27 Sept 2026
- paperEffective Approaches to Attention-based Neural Machine Translation, Luong, Pham and Manning · read 27 Sept 2026
- paperAttention Is All You Need, Vaswani et al. · read 27 Sept 2026
- docstorch.nn.functional.scaled_dot_product_attention, PyTorch · read 27 Sept 2026
- paperShow, Attend and Tell Neural Image Caption Generation with Visual Attention, Xu et al. · read 27 Sept 2026