Concepts

Sparse attention

3 min readintermediateUpdated 28 Sept 2026
1 · In one line

Sparse attention computes selected query-key connections instead of filling the entire attention matrix.

1 · What it is

Full self-attention grows quadratically with sequence length. Sparse attention computes selected query-key connections instead of filling the entire attention matrix.

A sparse mask defines which keys each query can attend to. The model evaluates attention only at those allowed pairs and mixes the values reached by the selected connections.

Longformer combines local windowed attention with task-motivated global attention. BigBird uses random, window and global attention components. Routing Transformer assigns queries and keys to clusters and attends within a cluster.

Hugging Face exposes block-sparse and original full attention modes for BigBird. Different sparse patterns make different connectivity trade-offs.

2 · Why it exists

Full self-attention grows quadratically with sequence length.

Pair countFull self-attention computes a weighting for each pair of sequence positions.
Long inputsQuadratic memory and computation make long sequences expensive.
Chosen routesSparse patterns compute only selected parts of the attention matrix.
3 · How it works

Apply a connectivity mask before attention is evaluated.

A sparse mask defines which keys each query can attend to.
  1. 1 · queriesStart with query, key and value vectors for the sequence.
  2. 2 · patternDefine a connectivity set for every output position.
  3. 3 · computeEvaluate attention only at the allowed query-key pairs.
  4. 4 · mixMix the values reached by those allowed connections.

Sparse attention is a family of patterns, not one universal mask.

4 · Where it's used
WhoWhat they askWhat it works with
Longformer“How can a long document combine nearby and task-level information?”Local windows plus task-motivated global attention
BigBird“Which sparse links should remain available?”Window, random and global connections
Routing Transformer“Which tokens have similar content?”Query and key clusters learned with online k-means
5 · What it solves, and what it doesn't
solves
  • Sparse factorizations can reduce the quadratic cost of full attention.
  • Longformer attention scales linearly with sequence length.
  • BigBird combines window, random and global attention components.
doesn't solve
  • A sparse pattern does not compute every pairwise query-key interaction.
  • Different sparse patterns make different connectivity trade-offs.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. paperGenerating Long Sequences with Sparse Transformers, Child et al. · read 28 Sept 2026
  2. paperLongformer The Long-Document Transformer, Beltagy, Peters and Cohan · read 28 Sept 2026
  3. paperBig Bird Transformers for Longer Sequences, Zaheer et al. · read 28 Sept 2026
  4. paperEfficient Content-Based Sparse Attention with Routing Transformers, Roy et al. · read 28 Sept 2026
  5. docsBigBird model documentation, Hugging Face · read 28 Sept 2026