Building with AI

Contextual retrieval

5 min readintermediateUpdated 28 Sept 2026
1 · In one line

Contextual retrieval has a model write a short note that places each chunk in its document, then indexes the note with the chunk so search can find it.

1 · What it is

Anthropic described contextual retrieval on 19 September 2024. Its goal is to improve the search step of RAG. In RAG (retrieval-augmented generation), a system first looks up passages in a collection of documents. Then it pastes them into the prompt, the text the model reads. Contextual retrieval changes the text that two different search methods see. The first is embedding search. An embedding model turns every chunk, a small piece of a document, into a vector: a list of numbers that captures what the chunk is about. The second is BM25, a scoring formula that ranks passages by how well their literal words match the query. BM25 also adjusts for how long a passage is, and it caps how much any single word can add, so stuffing a passage with one word stops paying off.

Both search indexes read the chunk text. So the fix is to improve that text, not the search. For each chunk, Anthropic prompted Claude 3 Haiku to write context specific to that chunk. In a code example, the note said the snippet belonged to the BasicConfigurator class in Log4cxx, a C++ logging library. Anthropic also tested pasting one general summary of the whole document onto its chunks. That helped very little in Anthropic’s experiments. Timing matters as well. HyDE, another search trick, has a language model invent a pretend answer when each question arrives. Contextual retrieval does its model work earlier, while building the index. So, unlike HyDE, searches are no slower.

Prompt caching, which lets the model reuse text it has already read, cuts that build-time cost. Every chunk from one document needs that full document in its prompt. The document is cached once, and each later request for that document’s chunks reads it back. Cache reads are billed at a tenth of the normal input price, and a cached entry lasts five minutes by default. Anthropic’s $1.02 figure for every million tokens of documents assumed 8,000-token documents split into 800-token chunks, each given about 100 tokens of context. The cookbook worked through a dataset of 737 chunks of code, and in that run 61.83% of input tokens came from cache.

The same cookbook measures quality with Pass@k. Pass@k asks a simple question. Does the one chunk that answers a test question, called the golden chunk, show up in the first k results? Among the top 10, that happened 87.15% of the time for plain RAG and 92.34% with contextual embeddings. A reranker is a model that takes the query plus candidate chunks and sorts them by relevance. The cookbook used Cohere’s rerank-english-v3.0 as its reranker. Adding one lifted the top-10 success rate to 95.26%.

2 · Why it exists

Cutting documents into chunks strips out the facts a search needs to find them.

Chunks lose their subjectA chunk can say revenue grew 3% without naming the company or the time period, so a search for that company may miss it.
Embeddings miss exact termsEmbedding search captures meaning but can miss exact matches, such as a query for the error code TS-999.
Missed chunksIn Anthropic's tests the baseline setup left 5.7% of relevant chunks outside the top 20 results.
3 · How it works

Follow one chunk from a company filing into the index and back out.

The model call happens once per chunk at ingestion. Every later search reads the stored context.
  1. 1 · splitThe knowledge base is broken into chunks, each one typically a few hundred tokens long at most.
  2. 2 · situateFor each chunk, a model reads the whole document plus the chunk and writes a short context, usually 50 to 100 tokens.
  3. 3 · prependThat context goes in front of the chunk before it is embedded and before the BM25 index is built.
  4. 4 · searchAt question time, embeddings and BM25 each find top chunks, and rank fusion merges and deduplicates the two lists.
  5. 5 · rerankOptionally, a reranker scores the top 150 chunks and passes the best 20 to the model.

The search itself does not change. Only the text being indexed changes, once, at ingestion.

4 · Where it's used
WhoWhat they askWhat it works with
Finance analyst“What was ACME's revenue growth in Q2 2023?”Filing chunks, each starting with a note on the company and quarter
Developer“What layout does BasicConfigurator use if none is given?”Code chunks, each starting with a note on the class and library
Support engineer“What does error TS-999 mean?”Help-article chunks matched on the exact code by contextual BM25
Legal team“Which past case discussed this clause?”Case chunks, each starting with a note on the case they come from
5 · What it solves, and what it doesn't
solves
  • Contextual embeddings cut the top-20 retrieval failure rate by 35%, from 5.7% to 3.7%.
  • Adding contextual BM25 cut it by 49%, to 2.9%.
  • With reranking on top, the cut was 67%, to 1.9%.
  • The context is written once at ingestion, not during every query.
doesn't solve
  • It adds one model call per chunk. With prompt caching, Anthropic put the cost at $1.02 for every million tokens of documents.
  • It does not pick chunk sizes for you. Chunk size, boundary and overlap still affect retrieval.
  • An embedding model with a fixed input limit may cut off the longer chunks and perform worse.
  • Each extra stage has a cost. Reranking adds a short delay and an extra API charge to every query.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. officialIntroducing Contextual Retrieval, Anthropic · read 29 Sept 2026
  2. officialContextual Retrieval Appendix II: full breakdown of experiment results, Anthropic · read 29 Sept 2026
  3. docsEnhancing RAG with contextual retrieval, Claude Cookbook, Anthropic · read 29 Sept 2026
  4. docsPrompt caching, Claude Docs, Anthropic · read 29 Sept 2026
  5. paperThe Probabilistic Relevance Framework: BM25 and Beyond, Robertson and Zaragoza, Foundations and Trends in Information Retrieval · read 29 Sept 2026
  6. paperPrecise Zero-Shot Dense Retrieval without Relevance Labels, Gao et al. · read 29 Sept 2026
  7. docsAn Overview of Cohere's Rerank Model, Cohere · read 29 Sept 2026