Chunking
Chunking splits a long document into smaller pieces that can be embedded, searched, or fitted into a model prompt.
Chunking cuts one long document into many smaller pieces of text. Anthropic describes the pieces as usually no more than a few hundred tokens. Each piece gets its own embedding, and a question fetches the pieces that are closest in meaning. Those pieces are then added to the prompt for the model.
The original RAG system split each Wikipedia article into separate 100-word chunks. That made 21 million passages. OpenAI’s file search cuts files into 800-token pieces by default. Neighbouring pieces share 400 tokens. Both numbers can be changed when you add a file.
A small piece can lose the facts that made it make sense. Anthropic’s fix is to write a short note about where each piece came from and put it in front of the piece before embedding. Together with a keyword index built the same way, that cut failed retrievals by 49% in their tests.
Search works on pieces, but careless cuts can hide the evidence a question needs.
Follow one handbook from pages to searchable pieces.
- 1 · measureA splitter sets a maximum piece size, which LangChain's recursive splitter counts in characters.
- 2 · cutIt tries separators in order, paragraph breaks first, then line breaks, then spaces, until every piece fits.
- 3 · overlapNeighbouring pieces can share some text, so a fact near a cut is less likely to be lost.
- 4 · storeEach piece is turned into an embedding and saved in a database searched by meaning.
Search returns pieces, not whole files, so the cuts decide what a question can find.
| Who | What they ask | What it works with |
|---|---|---|
| Support team | “What does the warranty cover?” | Overlapping sections of the current warranty |
| Developer portal | “How do I rotate a key?” | A procedure kept together under one heading |
| Legal team | “Which clause covers termination?” | Contract paragraphs with section metadata |
| Research assistant | “What evidence supports the result?” | Paper sections split around headings and paragraphs |
- Text longer than an embedding model's limit can still be embedded, one piece at a time.
- Splitting on paragraph breaks first keeps paragraphs whole for as long as the size allows.
- Overlap gives an idea that crosses a cut a better chance of surviving inside one piece.
- It does not pick the right size for you. Tools such as OpenAI's file search ship with default settings that you can change.
- A piece can still lack the names and dates needed to understand it.
- Good pieces do not make retrieval perfect. Anthropic still measured failed retrievals after adding context to each piece.
Sources used
This explainer is written in original language. The links below support its factual claims.
- docsRetrieval, OpenAI · read 27 Sept 2026
- officialIntroducing Contextual Retrieval, Anthropic · read 27 Sept 2026
- paperRetrieval-Augmented Generation for Knowledge-Intensive NLP Tasks, Lewis et al., NeurIPS 2020 · read 27 Sept 2026
- docsSplitting recursively, LangChain · read 27 Sept 2026
- docsEmbedding texts that are longer than the model's maximum context length, OpenAI Cookbook · read 27 Sept 2026