Building with AI

Chunking

3 min readbeginnerUpdated 28 Sept 2026
1 · In one line

Chunking splits a long document into smaller pieces that can be embedded, searched, or fitted into a model prompt.

1 · What it is

Chunking cuts one long document into many smaller pieces of text. Anthropic describes the pieces as usually no more than a few hundred tokens. Each piece gets its own embedding, and a question fetches the pieces that are closest in meaning. Those pieces are then added to the prompt for the model.

The original RAG system split each Wikipedia article into separate 100-word chunks. That made 21 million passages. OpenAI’s file search cuts files into 800-token pieces by default. Neighbouring pieces share 400 tokens. Both numbers can be changed when you add a file.

A small piece can lose the facts that made it make sense. Anthropic’s fix is to write a short note about where each piece came from and put it in front of the piece before embedding. Together with a keyword index built the same way, that cut failed retrievals by 49% in their tests.

2 · Why it exists

Search works on pieces, but careless cuts can hide the evidence a question needs.

Inputs have limitsAn embedding model has a maximum input length, counted in tokens. Longer text has to be cut short or split.
Cuts lose contextAnthropic's example is a piece that says the company's revenue grew 3%. On its own, it names neither the company nor the quarter.
Size is a trade-offHow big the pieces are, where the cuts fall and how much pieces overlap can all change what retrieval finds.
3 · How it works

Follow one handbook from pages to searchable pieces.

Paragraph breaks are tried first. Overlap keeps a little text on both sides of each cut.
  1. 1 · measureA splitter sets a maximum piece size, which LangChain's recursive splitter counts in characters.
  2. 2 · cutIt tries separators in order, paragraph breaks first, then line breaks, then spaces, until every piece fits.
  3. 3 · overlapNeighbouring pieces can share some text, so a fact near a cut is less likely to be lost.
  4. 4 · storeEach piece is turned into an embedding and saved in a database searched by meaning.

Search returns pieces, not whole files, so the cuts decide what a question can find.

4 · Where it's used
WhoWhat they askWhat it works with
Support team“What does the warranty cover?”Overlapping sections of the current warranty
Developer portal“How do I rotate a key?”A procedure kept together under one heading
Legal team“Which clause covers termination?”Contract paragraphs with section metadata
Research assistant“What evidence supports the result?”Paper sections split around headings and paragraphs
5 · What it solves, and what it doesn't
solves
  • Text longer than an embedding model's limit can still be embedded, one piece at a time.
  • Splitting on paragraph breaks first keeps paragraphs whole for as long as the size allows.
  • Overlap gives an idea that crosses a cut a better chance of surviving inside one piece.
doesn't solve
  • It does not pick the right size for you. Tools such as OpenAI's file search ship with default settings that you can change.
  • A piece can still lack the names and dates needed to understand it.
  • Good pieces do not make retrieval perfect. Anthropic still measured failed retrievals after adding context to each piece.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. docsRetrieval, OpenAI · read 27 Sept 2026
  2. officialIntroducing Contextual Retrieval, Anthropic · read 27 Sept 2026
  3. paperRetrieval-Augmented Generation for Knowledge-Intensive NLP Tasks, Lewis et al., NeurIPS 2020 · read 27 Sept 2026
  4. docsSplitting recursively, LangChain · read 27 Sept 2026
  5. docsEmbedding texts that are longer than the model's maximum context length, OpenAI Cookbook · read 27 Sept 2026