Rotary position embeddings
RoPE encodes absolute position with a rotation matrix and adds explicit relative-position dependence to self-attention.
RoPE encodes absolute position with a rotation matrix. It introduces an explicit relative-position dependency in self-attention.
The method rotates an affine-transformed embedding by angle multiples of its position index. The rotated query and key are then compared with an inner product.
LLaMA replaces absolute positional embeddings with RoPE. GPT-NeoX applies rotary embeddings to the first 25 percent of embedding-vector dimensions. RoFormer documentation describes rotations in two-dimensional space.
Llama implementations expose a RoPE base and optional scaling parameters. Implementations can apply RoPE to only part of the feature vector.
Attention needs position information inside token comparisons.
Apply positional rotations, then compare the results.
- 1 · projectSelf-attention transforms word embeddings into query, key and value representations.
- 2 · rotateRotate the affine-transformed embedding by angle multiples of its position index.
- 3 · compareCompute the inner product of the rotated query and key.
The rotations are position-dependent, while the final comparison is still a dot product.
| Who | What they ask | What it works with |
|---|---|---|
| RoFormer | “How can attention carry absolute and relative position information?” | Rotated query and key features |
| LLaMA | “Which position method should replace absolute position embeddings?” | Rotary position embeddings |
| GPT-NeoX | “How much of the embedding vector should receive rotary features?” | A configured rotary dimension fraction |
- It encodes absolute position with rotations.
- It introduces an explicit relative-position dependency in self-attention.
- It can replace learned absolute position embeddings in Transformer models.
- Llama implementations expose a RoPE base and optional scaling parameters.
- Implementations can apply RoPE to only part of the feature vector.
Sources used
This explainer is written in original language. The links below support its factual claims.
- paperRoFormer Enhanced Transformer with Rotary Position Embedding, Su et al. · read 28 Sept 2026
- paperLLaMA Open and Efficient Foundation Language Models, Touvron et al. · read 28 Sept 2026
- paperGPT-NeoX-20B An Open-Source Autoregressive Language Model, Black et al. · read 28 Sept 2026
- docsRoFormer model documentation, Hugging Face · read 28 Sept 2026
- docsLlama model documentation, Hugging Face · read 28 Sept 2026