Concepts

Long context

4 min readintermediateUpdated 28 Sept 2026
1 · In one line

Transformer-XL reuses hidden states from previous segments instead of computing them from scratch for each new segment.

1 · What it is

Long context is a design goal, not one mechanism. Longformer’s local-plus-global attention scales linearly with the input sequence. BigBird also uses sparse attention to reduce quadratic sequence-length dependence.

Transformer-XL instead uses segment-level recurrence and a positional encoding scheme. Position Interpolation rescales indices so a RoPE-based pretrained model can be adapted beyond its original context window.

Capacity and use are different. Lost in the Middle found that performance can degrade when relevant information appears in the middle of a long context. Extended-context models are not necessarily better at using their input context.

2 · Why it exists

Full self-attention becomes expensive as sequence length grows.

Quadratic attentionStandard self-attention scales quadratically with sequence length.
Position rangePosition Interpolation rescales position indices to fit a pretrained model's original window.
Retrieval qualityModels can perform worse when relevant information sits in the middle of a long input.
3 · How it works

One approach combines local windows with selected global attention.

Longformer combines local windowed attention with selected global attention so its attention computation scales linearly with sequence length.
  1. 1 · tokenizeRepresent the long document as a sequence of token positions.
  2. 2 · localGive each token a fixed-size attention window around its position.
  3. 3 · globalSelect task-relevant positions for global attention.
  4. 4 · combineUse local and global connections in the same attention mechanism.

The figure explains Longformer; recurrence and position interpolation extend context differently.

4 · Where it's used
WhoWhat they askWhat it works with
Document-system builder“Can one model pass include more of a long document?”Supported context window
Model researcher“Can attention cost grow more slowly with sequence length?”Sparse or linear-scaling attention
Evaluator“Does the model use evidence in every part of the window?”Position-sensitive long-context evaluation
5 · What it solves, and what it doesn't
solves
  • Sparse attention can reduce the quadratic sequence-length dependency to linear.
  • Segment recurrence can reuse hidden states from previous segments.
  • Position interpolation can extend RoPE-based pretrained models with fine-tuning.
doesn't solve
  • A larger input limit does not guarantee reliable use of middle-position evidence.
  • Extended-context models are not necessarily better at using their input context.
  • Full self-attention still scales quadratically with sequence length.