LlamaIndex
LlamaIndex is an open-source toolkit that brings your own files and data to a language model when you ask a question, so apps and agents can answer from them.
Language models learn from huge amounts of public text, but not from your own files. Your data may sit behind APIs, in databases, or inside PDFs and slide decks. LlamaIndex is an open-source toolkit for giving that data to the model at the moment a question is asked, instead of retraining the model.
Here is a concrete example. The official starter tutorial puts an essay by Paul Graham in a data folder. Then it asks what the author did in college. First, a reader (a loader) opens the file. Next, the index cuts it into small chunks called nodes. Then a query engine, the part that handles questions, finds the chunks that match. The model writes its answer from those chunks. This whole pattern is called retrieval-augmented generation, or RAG.
The same search can also become a tool. An agent is a program where the model decides, step by step, which tool to use next. In the same tutorial, one agent gets both a calculator and the essay search. Workflows link agents, data and other steps into bigger programs, where each step starts when something happens (an event).
A language model does not know your data, and cannot read all of it at once.
Follow one question about an essay through LlamaIndex.
- 1 · loadA data connector, also called a reader, pulls your files or records into Document objects of text and metadata.
- 2 · indexThe documents are split into chunks called nodes, and each node gets an embedding, a list of numbers standing for its meaning.
- 3 · storeThe index is usually saved, often in a vector database, so it does not have to be rebuilt.
- 4 · queryThe question becomes an embedding too, and the index finds the nodes whose numbers are most similar.
- 5 · answerA response synthesizer sends the question and those chunks to the model, which writes the reply.
LlamaIndex does the finding; the model does the writing.
| Who | What they ask | What it works with |
|---|---|---|
| Student | “What does my course reader say about the causes of the war?” | Chunks of the course PDFs |
| Support team | “Which plan includes phone support?” | Nodes from the pricing and help pages |
| Research team | “Which reports mention supply delays last year?” | An index of internal reports |
| Agent builder | “Can my agent search our docs and also do maths?” | A query engine wrapped as one of the agent's tools |
- Ready-made readers load data from many places, such as a local folder of files.
- More than 300 integration packages let you pick your own model, embedding and vector store providers.
- Only the relevant parts of your data go to the model with each question.
- A document search can be handed to an agent as one of its tools.
- It does not check answer quality for you; you still need to evaluate how accurate the replies are.
- By default the vector index lives only in memory, so lasting storage is a setup step.
- It is not a model itself; the starter tutorial calls a hosted OpenAI model with an API key.
- The company now focuses mainly on LlamaParse, its document-parsing platform, rather than the open-source toolkit.
Sources used
This explainer is written in original language. The links below support its factual claims.
- docsWelcome to LlamaIndex, LlamaIndex · read 28 Sept 2026
- docsHigh-Level Concepts, LlamaIndex · read 28 Sept 2026
- docsIntroduction to RAG, LlamaIndex · read 28 Sept 2026
- docsStarter Tutorial (Using OpenAI), LlamaIndex · read 28 Sept 2026
- docsUsing VectorStoreIndex, LlamaIndex · read 28 Sept 2026
- docsData Connectors, LlamaIndex · read 28 Sept 2026
- reporun-llama/llama_index, LlamaIndex on GitHub · read 28 Sept 2026