RAGBuilding with AI

Retrieval-Augmented Generation

4 min readbeginnerUpdated 28 Sept 2026
1 · In one line

RAG lets an AI model look up relevant documents first, then answer from what it found.

1 · What it is

A language model learns its facts during training and stores them inside itself. After training ends, that knowledge is fixed at a cut-off date. So its answers can be out of date or wrong. It also cannot easily show where an answer came from. When it does not know, it may make something up.

RAG adds a lookup step before the model answers. The system searches a set of trusted documents and hands the best matches to the model with the question. Picture an employee asking a company chatbot how much annual leave they get. The system pulls the leave policy from the handbook. The model then answers from that text, not from memory alone. Its answer can cite the policy, so the employee can open it and check.

The search works on meaning, not only on matching words. Each piece of text is stored as an embedding, a list of numbers that captures its meaning. Pieces with similar meanings end up close together, so a search for bad guy can find a passage about a villain.

For engineers: In the original research system, a neural retriever searched an index of Wikipedia passages. Wikipedia was cut into 100-word chunks, about 21 million in all. A model called BART wrote the answers. BART is a sequence-to-sequence model pretrained as a denoising autoencoder. Search scored each passage by the dot product of question and passage vectors, the method of Dense Passage Retrieval (DPR). The retriever and generator were trained together, without being told which document was right. Swapping in an older Wikipedia index changed what it knew, with no retraining. REALM had earlier used a retriever over Wikipedia during pre-training, fine-tuning and inference. The original RAG checkpoints were added to Hugging Face Transformers on 16 November 2020.

2 · Why it exists

On its own, a language model only knows what it learned in training.

Out of dateTraining stops at a cut-off date, so newer facts are missing.
No sourcesIt cannot easily show which fact led to an answer.
Makes things upWhen it lacks the answer, it may give false information.
3 · How it works

Follow one question from the documents to the answer.

When a question comes in, the search finds the best-matching passages and the model uses them as extra context.
  1. 1 · indexAhead of time, documents are split into small pieces, and each piece is stored as an embedding in a search index.
  2. 2 · findThe question is turned into an embedding too, and the closest pieces are pulled out.
  3. 3 · combineThose pieces are added to the prompt, next to the question.
  4. 4 · writeThe model writes its answer from the question and the retrieved text.

RAG gives the model something to read before it answers.

Swapping the documents was enough to change what the system knew.
4 · Where it's used
WhoWhat they askWhat it works with
Customer support“Can I return an opened item?”Passages from the current return policy
Law firm“Which contracts have a change-of-control clause?”Pieces of the contract archive
Engineering team“How do we rotate the API keys?”Passages from runbooks and wiki pages
HR“How much parental leave do we offer?”The current employee handbook
5 · What it solves, and what it doesn't
solves
  • You can update what it knows by changing the documents, without retraining the model.
  • Answers can point to their sources, so people can check them.
  • The looked-up text is plain text that a person can read.
  • In a writing test, people judged RAG more factual than a model without search in 42.7% of pairs, against 7.1% the other way.
doesn't solve
  • If the documents do not hold the answer, the search cannot find it.
  • The model may still answer from memory when nothing retrieved contains the answer.
  • It reduces made-up answers but does not end them.
  • The documents themselves can be wrong or biased.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. paperRetrieval-Augmented Generation for Knowledge-Intensive NLP Tasks, Lewis et al., NeurIPS 2020 · read 28 Sept 2026
  2. paperDense Passage Retrieval for Open-Domain Question Answering, Karpukhin et al. · read 28 Sept 2026
  3. paperBART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension, Lewis et al. · read 28 Sept 2026
  4. paperREALM: Retrieval-Augmented Language Model Pre-Training, Guu et al. · read 28 Sept 2026
  5. docsRAG, Hugging Face · read 28 Sept 2026
  6. docsWhat is RAG (Retrieval-Augmented Generation)?, Amazon Web Services · read 28 Sept 2026
  7. docsWhat is Retrieval-Augmented Generation (RAG)?, Google Cloud · read 28 Sept 2026