Positional encoding
Positional encoding adds token-order information to representations before attention reads them.
Self-attention alone does not supply a sequence order. A Transformer must inject information about token position.
The original Transformer adds a position vector to each token embedding. It uses sine and cosine functions of different frequencies. The position vector has the same dimension as the token embedding, so the two can be summed.
BERT uses absolute position embeddings. T5 uses a learned scalar that is added to an attention logit. ALiBi adds a static, non-learned bias to attention scores.
Absolute position embeddings and relative attention biases place the signal at different points in the computation. The choice of position scheme remains a model design decision.
Self-attention alone does not supply a sequence order.
Add a position vector before attention.
- 1 · tokensConvert each token into an embedding.
- 2 · positionsProduce one position vector with the same dimension as each token embedding.
- 3 · addAdd the token and position vectors element by element.
- 4 · attendSend the resulting sequence into the Transformer stack.
The position signal changes with the slot, even when the token identity stays the same.
| Who | What they ask | What it works with |
|---|---|---|
| Original Transformer | “Which sequence slot does this token occupy?” | Fixed sine and cosine position vectors |
| BERT | “Which absolute position belongs to this input token?” | Learned absolute position embeddings |
| T5 | “How far apart are this query and key?” | A learned scalar added to the attention logit |
- It supplies position information to an otherwise order-independent self-attention operation.
- Fixed sinusoidal encodings use sine and cosine functions of different frequencies.
- Relative schemes can represent the offset between a query and a key.
- The choice of position scheme remains a model design decision.
- Absolute position embeddings and relative attention biases place the signal at different points in the computation.
Sources used
This explainer is written in original language. The links below support its factual claims.
- paperAttention Is All You Need, Vaswani et al. · read 28 Sept 2026
- paperBERT Pre-training of Deep Bidirectional Transformers for Language Understanding, Devlin et al. · read 28 Sept 2026
- paperExploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer, Raffel et al. · read 28 Sept 2026
- paperTrain Short, Test Long Attention with Linear Biases Enables Input Length Extrapolation, Press, Smith and Lewis · read 28 Sept 2026
- docsBERT model documentation, Hugging Face · read 28 Sept 2026