Linear attention
Linear attention rewrites attention with feature maps so key-value summaries are formed before they are combined with queries.
Linear attention represents the similarity function with a kernel feature map. It first multiplies feature-mapped keys by values, then multiplies the result by feature-mapped queries.
For causal masking, the linear Transformer paper uses prefix sums. It also states that softmax attention has quadratic memory and time complexity in sequence length.
Performer uses FAVOR+ for softmax-kernel approximation. Efficient Attention switches the order of two matrix multiplications using associativity. Linearised self-attention can also be viewed as a fast weight memory system, while RetNet supports parallel, recurrent and chunkwise recurrent representations.
Standard dot-product attention forms pairwise interactions between all queries and keys.
Reorder one attention calculation around a key-value summary.
- 1 · mapApply feature maps to queries and keys.
- 2 · summarizeMultiply transformed keys by values to form a reusable summary.
- 3 · queryMultiply each transformed query by that summary.
- 4 · normalizeApply the corresponding feature-map normalization term.
Associativity allows the two matrix multiplications to switch order.
| Who | What they ask | What it works with |
|---|---|---|
| Sequence-model team | “Can attention avoid a full query-key matrix?” | Feature-mapped key-value summaries |
| Streaming engineer | “Can the summary be updated recurrently?” | Running key-value state |
| Kernel researcher | “Which feature map approximates the desired attention kernel?” | Kernel feature definitions |
- Linear Transformers express self-attention as a linear dot product of kernel feature maps.
- Performer uses FAVOR+ to approximate softmax attention kernels.
- Efficient Attention uses associativity to change the multiplication order.
- Performer approximates softmax and Gaussian kernels.
- The chosen feature map determines the similarity function.
- Causal linear attention requires a prefix-style accumulation.
Sources used
This explainer is written in original language. The links below support its factual claims.
- paperTransformers are RNNs, Katharopoulos et al. · read 28 Sept 2026
- paperRethinking Attention with Performers, Choromanski et al. · read 28 Sept 2026
- paperEfficient Attention, Shen et al. · read 28 Sept 2026
- paperLinear Transformers Are Secretly Fast Weight Programmers, Schlag, Irie and Schmidhuber · read 28 Sept 2026
- paperRetentive Network, Sun et al. · read 28 Sept 2026