Concepts

Rotary position embeddings

3 min readintermediateUpdated 28 Sept 2026
1 · In one line

RoPE encodes absolute position with a rotation matrix and adds explicit relative-position dependence to self-attention.

1 · What it is

RoPE encodes absolute position with a rotation matrix. It introduces an explicit relative-position dependency in self-attention.

The method rotates an affine-transformed embedding by angle multiples of its position index. The rotated query and key are then compared with an inner product.

LLaMA replaces absolute positional embeddings with RoPE. GPT-NeoX applies rotary embeddings to the first 25 percent of embedding-vector dimensions. RoFormer documentation describes rotations in two-dimensional space.

Llama implementations expose a RoPE base and optional scaling parameters. Implementations can apply RoPE to only part of the feature vector.

2 · Why it exists

Attention needs position information inside token comparisons.

Absolute positionRoPE encodes absolute position with a rotation matrix.
Relative relationThe query-key inner product carries an explicit relative-position dependency.
Direct operationThe method applies position through multiplication.
3 · How it works

Apply positional rotations, then compare the results.

RoPE rotates an affine-transformed embedding by angle multiples of its position index.
  1. 1 · projectSelf-attention transforms word embeddings into query, key and value representations.
  2. 2 · rotateRotate the affine-transformed embedding by angle multiples of its position index.
  3. 3 · compareCompute the inner product of the rotated query and key.

The rotations are position-dependent, while the final comparison is still a dot product.

4 · Where it's used
WhoWhat they askWhat it works with
RoFormer“How can attention carry absolute and relative position information?”Rotated query and key features
LLaMA“Which position method should replace absolute position embeddings?”Rotary position embeddings
GPT-NeoX“How much of the embedding vector should receive rotary features?”A configured rotary dimension fraction
5 · What it solves, and what it doesn't
solves
  • It encodes absolute position with rotations.
  • It introduces an explicit relative-position dependency in self-attention.
  • It can replace learned absolute position embeddings in Transformer models.
doesn't solve
  • Llama implementations expose a RoPE base and optional scaling parameters.
  • Implementations can apply RoPE to only part of the feature vector.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. paperRoFormer Enhanced Transformer with Rotary Position Embedding, Su et al. · read 28 Sept 2026
  2. paperLLaMA Open and Efficient Foundation Language Models, Touvron et al. · read 28 Sept 2026
  3. paperGPT-NeoX-20B An Open-Source Autoregressive Language Model, Black et al. · read 28 Sept 2026
  4. docsRoFormer model documentation, Hugging Face · read 28 Sept 2026
  5. docsLlama model documentation, Hugging Face · read 28 Sept 2026