Sparse attention
Sparse attention computes selected query-key connections instead of filling the entire attention matrix.
Full self-attention grows quadratically with sequence length. Sparse attention computes selected query-key connections instead of filling the entire attention matrix.
A sparse mask defines which keys each query can attend to. The model evaluates attention only at those allowed pairs and mixes the values reached by the selected connections.
Longformer combines local windowed attention with task-motivated global attention. BigBird uses random, window and global attention components. Routing Transformer assigns queries and keys to clusters and attends within a cluster.
Hugging Face exposes block-sparse and original full attention modes for BigBird. Different sparse patterns make different connectivity trade-offs.
Full self-attention grows quadratically with sequence length.
Apply a connectivity mask before attention is evaluated.
- 1 · queriesStart with query, key and value vectors for the sequence.
- 2 · patternDefine a connectivity set for every output position.
- 3 · computeEvaluate attention only at the allowed query-key pairs.
- 4 · mixMix the values reached by those allowed connections.
Sparse attention is a family of patterns, not one universal mask.
| Who | What they ask | What it works with |
|---|---|---|
| Longformer | “How can a long document combine nearby and task-level information?” | Local windows plus task-motivated global attention |
| BigBird | “Which sparse links should remain available?” | Window, random and global connections |
| Routing Transformer | “Which tokens have similar content?” | Query and key clusters learned with online k-means |
- Sparse factorizations can reduce the quadratic cost of full attention.
- Longformer attention scales linearly with sequence length.
- BigBird combines window, random and global attention components.
- A sparse pattern does not compute every pairwise query-key interaction.
- Different sparse patterns make different connectivity trade-offs.
Sources used
This explainer is written in original language. The links below support its factual claims.
- paperGenerating Long Sequences with Sparse Transformers, Child et al. · read 28 Sept 2026
- paperLongformer The Long-Document Transformer, Beltagy, Peters and Cohan · read 28 Sept 2026
- paperBig Bird Transformers for Longer Sequences, Zaheer et al. · read 28 Sept 2026
- paperEfficient Content-Based Sparse Attention with Routing Transformers, Roy et al. · read 28 Sept 2026
- docsBigBird model documentation, Hugging Face · read 28 Sept 2026