Building with AI

Context engineering

4 min readintermediateUpdated 28 Sept 2026
1 · In one line

Context engineering is choosing, trimming and updating the tokens a model sees at each step, so an AI agent works from a small, useful context.

1 · What it is

Prompt engineering is the craft of wording a model’s instructions. Context engineering handles everything else the model reads each time it is called. That includes tool descriptions, earlier turns of the chat, outside data and whatever its tools sent back. Tools let an agent (an AI that works through a task step by step) act on the world. It might read a file or run a search, and each tool hands back a tool result. An agent calls the model many times, so it decides what to pack again before every call.

Size is counted in tokens, which are small chunks of text. For Gemini models, 100 tokens is about 60 to 80 English words. Claude’s context editing feature shows trimming with real numbers. In the documented setup, clearing starts once a request passes 30,000 input tokens. The API (the service programs use to talk to Claude) then deletes tool results, oldest first. It keeps the 3 most recent tool calls with their results, and each round must clear at least 5,000 tokens. Each deleted result leaves a short placeholder, so Claude can tell something used to be there. Compaction is the heavier option. Instead of dropping single results, the server replaces earlier turns with a summary. Anthropic’s API offers it as a built-in option.

Where material sits in the window matters too. A study called Lost in the Middle tested this with long inputs. Models often did best when a fact sat near the start or the end. They did much worse when it sat in the middle. Google gives Gemini users two tips that fit this. Ask your question after the long material. And leave out tokens the model does not need, even though many Gemini models accept a million tokens or more.

2 · Why it exists

A context window is a limited budget, and filling it carelessly makes models worse.

Everything countsOne window has to hold the system prompt (the standing instructions), every message, every tool result, the tool descriptions and the model's own reply.
More can mean worsePack in more tokens and the model recalls what they say less reliably, an effect Anthropic calls context rot.
Agents pile upTools an agent runs hand back more text, such as a file's contents or a list of search hits. Over a long task that heap keeps growing, and any part of it might matter later.
3 · How it works

Follow one turn of a coding agent as its window is rebuilt.

Old tool results go first: once the model has processed them, they are rarely needed again.
  1. 1 · gatherCollect everything that could go in, from instructions and tool descriptions to outside data and the chat so far.
  2. 2 · fetchHold pointers such as a file path, and open the file itself only on the step that needs it.
  3. 3 · trimPast a size limit you choose, drop old tool results, oldest first, and leave a marker in each gap.
  4. 4 · compactBefore the window runs out, have the model write up the work so far, and let that write-up open a new, nearly empty window.
  5. 5 · rememberSave progress to a notes file that lives outside the window, and load it back when it becomes useful.

Default to less: every token in the window should earn its place for this step.

4 · Where it's used
WhoWhat they askWhat it works with
Coding agent“Why does the build fail after this refactor?”Recent file reads and searches, with older tool results cleared
Lead agent“What did each helper find?”Condensed summaries returned by subagents
Long-running chat“Can we keep going past the window limit?”A summary of the earlier conversation
Game-playing agent“Where was I in my training?”Its own notes, read back after a context reset
5 · What it solves, and what it doesn't
solves
  • Chats and tasks can outlast a single window, because older turns get folded into a summary.
  • A lead agent can send helper agents (sub-agents) off to explore. Each may spend tens of thousands of tokens, then report back in about 1,000 to 2,000.
  • Over thousands of moves in a Pokémon game, Anthropic's agent kept exact counts by writing notes to itself.
  • Answers improve when the question comes after long documents, as Google advises for Gemini.
doesn't solve
  • It does not make the window bigger. Prompt caching (reusing an already processed opening of a prompt) changes what you pay, but those tokens still fill space.
  • Summaries lose things. A detail that looked minor when the summary was written may be the one the agent needs later.
  • Just-in-time loading has a cost. Each lookup takes time that data prepared in advance would not.
  • A larger window is no escape. In RULER, a benchmark for long inputs, only half of the models advertising windows of 32,000 tokens or more still did well at that length.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. officialEffective context engineering for AI agents, Anthropic · read 28 Sept 2026
  2. docsContext windows, Anthropic · read 28 Sept 2026
  3. docsContext editing, Anthropic · read 28 Sept 2026
  4. docsLong context, Google · read 28 Sept 2026
  5. paperLost in the Middle: How Language Models Use Long Contexts, Liu et al., TACL 2023 · read 28 Sept 2026
  6. paperRULER: What's the Real Context Size of Your Long-Context Language Models?, Hsieh et al., COLM 2024 · read 28 Sept 2026
  7. docsUnderstand and count tokens, Google · read 28 Sept 2026