Concepts

Self-attention

3 min readbeginnerUpdated 28 Sept 2026
1 · In one line

Self-attention updates each position by comparing it with positions in the same sequence and mixing their information.

1 · What it is

Self-attention is attention whose queries, keys and values are derived from the same sequence. Self-attention takes vectors, not raw words, as its inputs. A learned projection gives each position its query, key and value vectors.

For one output position, the query is compared with the permitted keys. Softmax changes the scores into weights on the values. The calculation repeats for every position.

Encoder self-attention lets each position attend to all positions in the previous encoder layer. BERT uses bidirectional self-attention. Decoder masking prevents positions from attending to subsequent positions. GPT-3’s decoder uses an attention mask in its self-attention pattern. That restriction is implemented with a mask before softmax.

2 · Why it exists

A token's useful meaning often depends on other positions in the same sequence.

Distant contextSelf-attention connects positions without making information travel through every position in between.
Same sequenceIts queries, keys and values all come from the previous layer's representations of one sequence.
Future leakageA decoder must mask later positions so a prediction cannot read tokens it is supposed to predict.
3 · How it works

Follow the position for bank as it reads its own sentence.

Self-attention takes its queries, keys and values from the same sequence.
  1. 1 · projectEach input representation is projected into a query, a key and a value.
  2. 2 · compareOne position's query is compared with the keys from permitted positions in the sequence.
  3. 3 · maskIn an autoregressive decoder, scores for later positions are blocked before softmax.
  4. 4 · mixThe normalized weights mix the permitted value vectors into an updated representation for that position.

Self-attention describes where Q, K and V come from: the same sequence, not necessarily the same vector.

4 · Where it's used
WhoWhat they askWhat it works with
Language encoder“Which words help represent this word?”All permitted token positions in the input
Text generator“Which earlier tokens matter for the next-token state?”The generated prefix, with later positions masked
Vision transformer“Which image patches help represent this patch?”Patch representations in the same image
5 · What it solves, and what it doesn't
solves
  • Every position can directly draw information from other permitted positions in one layer.
  • A self-attention layer connects all positions with a constant number of sequential operations.
  • A causal mask can enforce the left-to-right information boundary used by an autoregressive decoder.
doesn't solve
  • Self-attention needs position information because attention alone does not encode sequence order.
  • Full self-attention computes a score for each pair of positions, so its per-layer complexity grows quadratically with sequence length.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. paperAttention Is All You Need, Vaswani et al. · read 27 Sept 2026
  2. paperBERT Pre-training of Deep Bidirectional Transformers for Language Understanding, Devlin et al. · read 27 Sept 2026
  3. docstorch.nn.functional.scaled_dot_product_attention, PyTorch · read 27 Sept 2026
  4. paperLanguage Models are Few-Shot Learners, Brown et al. · read 27 Sept 2026
  5. paperLongformer The Long-Document Transformer, Beltagy, Peters and Cohan · read 27 Sept 2026