Concepts

Linear attention

3 min readadvancedUpdated 28 Sept 2026
1 · In one line

Linear attention rewrites attention with feature maps so key-value summaries are formed before they are combined with queries.

1 · What it is

Linear attention represents the similarity function with a kernel feature map. It first multiplies feature-mapped keys by values, then multiplies the result by feature-mapped queries.

For causal masking, the linear Transformer paper uses prefix sums. It also states that softmax attention has quadratic memory and time complexity in sequence length.

Performer uses FAVOR+ for softmax-kernel approximation. Efficient Attention switches the order of two matrix multiplications using associativity. Linearised self-attention can also be viewed as a fast weight memory system, while RetNet supports parallel, recurrent and chunkwise recurrent representations.

2 · Why it exists

Standard dot-product attention forms pairwise interactions between all queries and keys.

Pairwise matrixStandard attention computes a query-key product before multiplying by values.
Sequence costThe linear Transformer paper states that softmax attention has quadratic memory and time complexity in sequence length.
Different kernelLinear attention replaces softmax similarity with a kernel represented by feature maps.
3 · How it works

Reorder one attention calculation around a key-value summary.

Linear attention key-value summary Transformed keys and values are multiplied first to create a reusable summary. A transformed query then reads that summary and applies a normalization term to produce its output.
Associativity lets the model combine keys with values first, then apply each transformed query to that summary.
  1. 1 · mapApply feature maps to queries and keys.
  2. 2 · summarizeMultiply transformed keys by values to form a reusable summary.
  3. 3 · queryMultiply each transformed query by that summary.
  4. 4 · normalizeApply the corresponding feature-map normalization term.

Associativity allows the two matrix multiplications to switch order.

4 · Where it's used
WhoWhat they askWhat it works with
Sequence-model team“Can attention avoid a full query-key matrix?”Feature-mapped key-value summaries
Streaming engineer“Can the summary be updated recurrently?”Running key-value state
Kernel researcher“Which feature map approximates the desired attention kernel?”Kernel feature definitions
5 · What it solves, and what it doesn't
solves
  • Linear Transformers express self-attention as a linear dot product of kernel feature maps.
  • Performer uses FAVOR+ to approximate softmax attention kernels.
  • Efficient Attention uses associativity to change the multiplication order.
doesn't solve
  • Performer approximates softmax and Gaussian kernels.
  • The chosen feature map determines the similarity function.
  • Causal linear attention requires a prefix-style accumulation.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. paperTransformers are RNNs, Katharopoulos et al. · read 28 Sept 2026
  2. paperRethinking Attention with Performers, Choromanski et al. · read 28 Sept 2026
  3. paperEfficient Attention, Shen et al. · read 28 Sept 2026
  4. paperLinear Transformers Are Secretly Fast Weight Programmers, Schlag, Irie and Schmidhuber · read 28 Sept 2026
  5. paperRetentive Network, Sun et al. · read 28 Sept 2026