Building with AI

LlamaIndex

4 min readintermediateUpdated 28 Sept 2026
1 · In one line

LlamaIndex is an open-source toolkit that brings your own files and data to a language model when you ask a question, so apps and agents can answer from them.

1 · What it is

Language models learn from huge amounts of public text, but not from your own files. Your data may sit behind APIs, in databases, or inside PDFs and slide decks. LlamaIndex is an open-source toolkit for giving that data to the model at the moment a question is asked, instead of retraining the model.

Here is a concrete example. The official starter tutorial puts an essay by Paul Graham in a data folder. Then it asks what the author did in college. First, a reader (a loader) opens the file. Next, the index cuts it into small chunks called nodes. Then a query engine, the part that handles questions, finds the chunks that match. The model writes its answer from those chunks. This whole pattern is called retrieval-augmented generation, or RAG.

The same search can also become a tool. An agent is a program where the model decides, step by step, which tool to use next. In the same tutorial, one agent gets both a calculator and the essay search. Workflows link agents, data and other steps into bigger programs, where each step starts when something happens (an event).

2 · Why it exists

A language model does not know your data, and cannot read all of it at once.

Not trained on youModels learn from huge amounts of public text, not from your notes, files or company records.
Data is scatteredUseful data often lives behind APIs, inside SQL databases, or in PDFs and slide decks.
Too much to sendSending everything with every question is wasteful, so only the relevant parts should go to the model.
3 · How it works

Follow one question about an essay through LlamaIndex.

LlamaIndex: documents are loaded, split into nodes and indexed once; each question is matched against the index and the closest nodes go to the model ONCE, AHEAD OF TIME EVERY QUESTION Data folder Reader Split into nodes Vector store index Question Query engine Closest nodes Model answers an essay as a .txt file makes Document objects chunks with metadata one embedding per node What did the author do in college? question to embedding most similar numbers response synthesizer + language model stored index WHAT THE MODEL ACTUALLY READS The question plus the few retrieved chunks, not the whole essay. The model writes its reply from that small bundle.
The index is built once. Each question then fetches only the few nodes that match it.
  1. 1 · loadA data connector, also called a reader, pulls your files or records into Document objects of text and metadata.
  2. 2 · indexThe documents are split into chunks called nodes, and each node gets an embedding, a list of numbers standing for its meaning.
  3. 3 · storeThe index is usually saved, often in a vector database, so it does not have to be rebuilt.
  4. 4 · queryThe question becomes an embedding too, and the index finds the nodes whose numbers are most similar.
  5. 5 · answerA response synthesizer sends the question and those chunks to the model, which writes the reply.

LlamaIndex does the finding; the model does the writing.

4 · Where it's used
WhoWhat they askWhat it works with
Student“What does my course reader say about the causes of the war?”Chunks of the course PDFs
Support team“Which plan includes phone support?”Nodes from the pricing and help pages
Research team“Which reports mention supply delays last year?”An index of internal reports
Agent builder“Can my agent search our docs and also do maths?”A query engine wrapped as one of the agent's tools
5 · What it solves, and what it doesn't
solves
  • Ready-made readers load data from many places, such as a local folder of files.
  • More than 300 integration packages let you pick your own model, embedding and vector store providers.
  • Only the relevant parts of your data go to the model with each question.
  • A document search can be handed to an agent as one of its tools.
doesn't solve
  • It does not check answer quality for you; you still need to evaluate how accurate the replies are.
  • By default the vector index lives only in memory, so lasting storage is a setup step.
  • It is not a model itself; the starter tutorial calls a hosted OpenAI model with an API key.
  • The company now focuses mainly on LlamaParse, its document-parsing platform, rather than the open-source toolkit.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. docsWelcome to LlamaIndex, LlamaIndex · read 28 Sept 2026
  2. docsHigh-Level Concepts, LlamaIndex · read 28 Sept 2026
  3. docsIntroduction to RAG, LlamaIndex · read 28 Sept 2026
  4. docsStarter Tutorial (Using OpenAI), LlamaIndex · read 28 Sept 2026
  5. docsUsing VectorStoreIndex, LlamaIndex · read 28 Sept 2026
  6. docsData Connectors, LlamaIndex · read 28 Sept 2026
  7. reporun-llama/llama_index, LlamaIndex on GitHub · read 28 Sept 2026