Self-attention
Self-attention updates each position by comparing it with positions in the same sequence and mixing their information.
Self-attention is attention whose queries, keys and values are derived from the same sequence. Self-attention takes vectors, not raw words, as its inputs. A learned projection gives each position its query, key and value vectors.
For one output position, the query is compared with the permitted keys. Softmax changes the scores into weights on the values. The calculation repeats for every position.
Encoder self-attention lets each position attend to all positions in the previous encoder layer. BERT uses bidirectional self-attention. Decoder masking prevents positions from attending to subsequent positions. GPT-3’s decoder uses an attention mask in its self-attention pattern. That restriction is implemented with a mask before softmax.
A token's useful meaning often depends on other positions in the same sequence.
Follow the position for bank as it reads its own sentence.
- 1 · projectEach input representation is projected into a query, a key and a value.
- 2 · compareOne position's query is compared with the keys from permitted positions in the sequence.
- 3 · maskIn an autoregressive decoder, scores for later positions are blocked before softmax.
- 4 · mixThe normalized weights mix the permitted value vectors into an updated representation for that position.
Self-attention describes where Q, K and V come from: the same sequence, not necessarily the same vector.
| Who | What they ask | What it works with |
|---|---|---|
| Language encoder | “Which words help represent this word?” | All permitted token positions in the input |
| Text generator | “Which earlier tokens matter for the next-token state?” | The generated prefix, with later positions masked |
| Vision transformer | “Which image patches help represent this patch?” | Patch representations in the same image |
- Every position can directly draw information from other permitted positions in one layer.
- A self-attention layer connects all positions with a constant number of sequential operations.
- A causal mask can enforce the left-to-right information boundary used by an autoregressive decoder.
- Self-attention needs position information because attention alone does not encode sequence order.
- Full self-attention computes a score for each pair of positions, so its per-layer complexity grows quadratically with sequence length.
Sources used
This explainer is written in original language. The links below support its factual claims.
- paperAttention Is All You Need, Vaswani et al. · read 27 Sept 2026
- paperBERT Pre-training of Deep Bidirectional Transformers for Language Understanding, Devlin et al. · read 27 Sept 2026
- docstorch.nn.functional.scaled_dot_product_attention, PyTorch · read 27 Sept 2026
- paperLanguage Models are Few-Shot Learners, Brown et al. · read 27 Sept 2026
- paperLongformer The Long-Document Transformer, Beltagy, Peters and Cohan · read 27 Sept 2026