Context compaction
Context compaction swaps the older part of a long AI conversation for a short summary, so the chat can keep going inside the model's memory limit.
A language model reads everything in its context window, a kind of working memory for one response. That window has a size limit. More text is not always better, either. As the window fills, the model gets worse at recalling what is inside it. Long chats hit both problems. So do agents, AI helpers that take actions by calling tools, when they call many of them.
Anthropic calls compaction the usual first step when managing an agent’s context. It works like writing a short recap of a long group chat. When a conversation nears the limit, a model summarises what has happened so far. A fresh window then starts with that summary in place of the old turns. Anthropic gives Claude Code as an example. After a long session, its summary keeps design decisions, open bugs and implementation details, and drops repeated tool output. Claude Code then carries on with the summary plus the five files it used most recently.
You can start it yourself or let it run on its own. In Claude Code you can type /compact with instructions about what to keep. It also compacts automatically when it nears the limit. If you want a clean start rather than continuity, /clear costs nothing.
For engineers: Anthropic’s API can compact when you ask it to, or automatically once the token count (tokens are the small chunks of text a model reads) passes a set level. OpenAI’s Responses API can compact on the server when a set token count is crossed. OpenAI also offers a separate compact endpoint that returns a new, smaller context window. Google’s Agent Development Kit can compact by token count or after a fixed number of turns. A lighter form of compaction simply clears old tool results.
Long conversations run into three problems.
Follow one long agent session as it gets compacted.
- 1 · growThe conversation and its tool results keep adding tokens, the small pieces of text a model reads.
- 2 · triggerWhen the token count crosses a set threshold, or when you ask, compaction starts.
- 3 · summariseA model reads the older history and writes a summary of what later steps will need.
- 4 · continueThe summary replaces the older turns, recent turns can stay word for word, and work carries on.
Compaction trades old detail for room to keep going.
| Who | What they ask | What it works with |
|---|---|---|
| Coding agent | “Keep working on this bug fix after hours of edits” | A long history of file reads, test runs and decisions |
| Support chatbot | “What did I say about my order earlier” | A long chat with one customer |
| Research assistant | “Add two more sources to the report” | Many earlier search results and drafts |
| Trip planner | “Stretch the trip plan by another two days” | The plan built over many turns |
- Keeps a chat or agent job going after the context window fills up.
- Keeps the active context small, which helps answer quality.
- Sends fewer tokens on later requests, which can lower cost and waiting time.
- Can keep the most recent turns exactly as they were.
- A summary that cuts too much can lose a detail that turns out to matter later.
- Compacting a very long history is itself a large request.
- Memory beyond the context window is a separate technique: the agent writes notes it keeps outside the window.
- Some compacted forms are encrypted and cannot be read by people.
Sources used
This explainer is written in original language. The links below support its factual claims.
- docsCompaction overview, Anthropic · read 28 Sept 2026
- officialEffective context engineering for AI agents, Anthropic · read 28 Sept 2026
- docsManage costs effectively, Anthropic (Claude Code docs) · read 28 Sept 2026
- docsContext windows, Anthropic · read 28 Sept 2026
- docsCompaction, OpenAI · read 28 Sept 2026
- docsCompress agent context for performance, Google Agent Development Kit · read 28 Sept 2026