Long context
Transformer-XL reuses hidden states from previous segments instead of computing them from scratch for each new segment.
Long context is a design goal, not one mechanism. Longformer’s local-plus-global attention scales linearly with the input sequence. BigBird also uses sparse attention to reduce quadratic sequence-length dependence.
Transformer-XL instead uses segment-level recurrence and a positional encoding scheme. Position Interpolation rescales indices so a RoPE-based pretrained model can be adapted beyond its original context window.
Capacity and use are different. Lost in the Middle found that performance can degrade when relevant information appears in the middle of a long context. Extended-context models are not necessarily better at using their input context.
Full self-attention becomes expensive as sequence length grows.
One approach combines local windows with selected global attention.
- 1 · tokenizeRepresent the long document as a sequence of token positions.
- 2 · localGive each token a fixed-size attention window around its position.
- 3 · globalSelect task-relevant positions for global attention.
- 4 · combineUse local and global connections in the same attention mechanism.
The figure explains Longformer; recurrence and position interpolation extend context differently.
| Who | What they ask | What it works with |
|---|---|---|
| Document-system builder | “Can one model pass include more of a long document?” | Supported context window |
| Model researcher | “Can attention cost grow more slowly with sequence length?” | Sparse or linear-scaling attention |
| Evaluator | “Does the model use evidence in every part of the window?” | Position-sensitive long-context evaluation |
- Sparse attention can reduce the quadratic sequence-length dependency to linear.
- Segment recurrence can reuse hidden states from previous segments.
- Position interpolation can extend RoPE-based pretrained models with fine-tuning.
- A larger input limit does not guarantee reliable use of middle-position evidence.
- Extended-context models are not necessarily better at using their input context.
- Full self-attention still scales quadratically with sequence length.
Sources used
This explainer is written in original language. The links below support its factual claims.
- paperLongformer - The Long-Document Transformer, arXiv · read 28 Sept 2026
- paperBig Bird - Transformers for Longer Sequences, NeurIPS / arXiv · read 28 Sept 2026
- paperTransformer-XL - Attentive Language Models Beyond a Fixed-Length Context, ACL / arXiv · read 28 Sept 2026
- paperLost in the Middle - How Language Models Use Long Contexts, TACL / arXiv · read 28 Sept 2026
- paperExtending Context Window of Large Language Models via Positional Interpolation, arXiv · read 28 Sept 2026