Concepts

Context window

3 min readbeginnerUpdated 28 Sept 2026
1 · In one line

A context window is all the text a language model can reference while generating a response, including the response itself.

1 · What it is

The context window states the material available for one model response. The generated response also counts toward the context window. Context-window sizes vary by model.

GPT-2 increased its context size from 512 to 1,024 tokens. GPT-3 trained on sequences filling a 2,048-token context window. BERT training used combined sequences no longer than 512 tokens. Longformer added position embeddings supporting positions up to 4,096.

2 · Why it exists

A language model can reference only the material inside its context window.

Inputs add upSystem prompts, messages, tool results, images, documents and tool definitions count toward the window.
Output countsThe generated response also counts toward the context window.
Capacity differsContext-window sizes are listed by model.
3 · How it works

Account for one request inside a fixed illustrative window.

The 16 slots are illustrative, not a real model limit; the highlighted capacity boundary is step 3.
  1. 1 · assembleThe request includes system instructions, messages, tool results, documents and tool definitions.
  2. 2 · reserveThe generated output also uses space inside the context window.
  3. 3 · fitIf the input alone exceeds the supported window, the Claude API returns an invalid request error.
  4. 4 · generateThe model generates while referencing the text included in that window.

Context-window sizes vary by model.

4 · Where it's used
WhoWhat they askWhat it works with
Application developer“Will this request and its answer fit together?”The selected model's documented context window
Agent engineer“Which tool results are consuming request capacity?”Every message and tool result sent in the request
Document assistant“Can this entire document be included at once?”Its token count and the selected model limit
Model researcher“How does longer input change attention cost?”The architecture's attention pattern and sequence length
5 · What it solves, and what it doesn't
solves
  • The context window states the material available for one model response.
  • Everything in the request counts toward the context window.
  • The Claude API rejects an input that already exceeds the model's context window.
doesn't solve
  • A large context window does not mean every possible input will fit.
  • A larger context does not guarantee accuracy and recall across all included text.
  • Extending sequence length does not remove the computational cost of attention.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. docsContext windows, Anthropic · read 28 Sept 2026
  2. paperLanguage Models are Unsupervised Multitask Learners, OpenAI · read 28 Sept 2026
  3. paperLanguage Models are Few-Shot Learners, Brown and colleagues · read 28 Sept 2026
  4. paperBERT Pre-training of Deep Bidirectional Transformers for Language Understanding, Devlin and colleagues · read 28 Sept 2026
  5. paperLongformer: The Long-Document Transformer, Beltagy, Peters and Cohan · read 28 Sept 2026