Concepts

Positional encoding

3 min readintermediateUpdated 28 Sept 2026
1 · In one line

Positional encoding adds token-order information to representations before attention reads them.

1 · What it is

Self-attention alone does not supply a sequence order. A Transformer must inject information about token position.

The original Transformer adds a position vector to each token embedding. It uses sine and cosine functions of different frequencies. The position vector has the same dimension as the token embedding, so the two can be summed.

BERT uses absolute position embeddings. T5 uses a learned scalar that is added to an attention logit. ALiBi adds a static, non-learned bias to attention scores.

Absolute position embeddings and relative attention biases place the signal at different points in the computation. The choice of position scheme remains a model design decision.

2 · Why it exists

Self-attention alone does not supply a sequence order.

Order signalA Transformer must inject information about token position.
Compatible shapeThe original Transformer made position and token vectors the same size so they could be summed.
More than one designPositional encodings can be learned or fixed.
3 · How it works

Add a position vector before attention.

The original Transformer adds a position vector to each token embedding.
  1. 1 · tokensConvert each token into an embedding.
  2. 2 · positionsProduce one position vector with the same dimension as each token embedding.
  3. 3 · addAdd the token and position vectors element by element.
  4. 4 · attendSend the resulting sequence into the Transformer stack.

The position signal changes with the slot, even when the token identity stays the same.

4 · Where it's used
WhoWhat they askWhat it works with
Original Transformer“Which sequence slot does this token occupy?”Fixed sine and cosine position vectors
BERT“Which absolute position belongs to this input token?”Learned absolute position embeddings
T5“How far apart are this query and key?”A learned scalar added to the attention logit
5 · What it solves, and what it doesn't
solves
  • It supplies position information to an otherwise order-independent self-attention operation.
  • Fixed sinusoidal encodings use sine and cosine functions of different frequencies.
  • Relative schemes can represent the offset between a query and a key.
doesn't solve
  • The choice of position scheme remains a model design decision.
  • Absolute position embeddings and relative attention biases place the signal at different points in the computation.