Building with AI

Haystack

4 min readintermediateUpdated 28 Sept 2026
1 · In one line

An open-source Python toolkit from deepset that builds AI apps (search, RAG, agents) by linking small parts into a pipeline.

1 · What it is

Haystack, built by the company deepset, is an open-source Python framework for building apps on top of large language models (LLMs). Its code is free to use under the Apache-2.0 licence, a common open-source licence. People use it for search tools, for RAG apps that answer from their own documents, and for agents that call tools.

The main idea is the pipeline. A pipeline is a chain of small parts called components. Each component does one job, such as turning text into an embedding (a list of numbers that captures meaning), fetching documents or calling a model. You add the components, then connect one part’s named output to the next part’s named input. Haystack checks that each output and input hold the same kind of data. It does this when you connect them, before anything runs.

Picture a school library that wants a help bot. Someone asks, “Is the library open on Saturdays?” An embedder turns the question into numbers. A retriever uses them to fetch the best-matching pages from a document store, the database that holds the library’s documents. A prompt builder puts those pages next to the question. The model then writes an answer from them.

Pipelines do not have to be a straight line. They can split into branches, for example to send PDFs and web pages to different converters. They can also loop, so a checker can send a bad answer back for another try, up to a limit you set. For open-ended tasks there is a built-in Agent component, which lets the model call tools until it reaches a stopping point. A whole pipeline can even become one of its tools. Pipelines can be saved as YAML files, a plain text format for settings. A companion tool called Hayhooks can serve them as web APIs, so other apps can send them requests over the internet, or as MCP servers that AI assistants can connect to.

2 · Why it exists

An AI app that answers from your own files needs several parts working together.

Many moving partsStoring documents, finding the right ones, building a prompt and calling a model are separate jobs.
Parts must fitOne part's output has to be the kind of data the next part expects.
Models keep changingTeams want to swap one model or database for another without rewriting the whole app.
3 · How it works

Follow one question through a Haystack RAG pipeline.

This is the query pipeline from the Haystack guide to creating pipelines.
  1. 1 · addYou add components to a pipeline, each one doing a single task such as embedding or retrieving.
  2. 2 · connectYou connect one component's named output to the next one's input, and Haystack checks the types match.
  3. 3 · retrieveWhen the pipeline runs, the retriever fetches matching documents from the document store.
  4. 4 · generateThe prompt builder adds those documents to the prompt, and the model writes the answer.

In Haystack the pipeline decides which part passes what to which.

4 · Where it's used
WhoWhat they askWhat it works with
Support team“Can our help bot answer from the product manuals?”A RAG pipeline over the manuals
Operations team“Can one pipeline read PDFs, web pages and text files?”Branches that send each file type to its own converter
Developer“Can the model check its own output and try again?”A loop between a generator and a validator
Product team“Can the assistant look things up and call our tools?”The built-in Agent component with tools
5 · What it solves, and what it doesn't
solves
  • It gives ready-made components for jobs like retrieval, indexing, tool calling and memory.
  • It checks that connected parts fit before the pipeline runs.
  • It works with many model makers, so you can swap models without rewriting the system.
  • A pipeline can be saved as YAML and loaded again.
doesn't solve
  • You still need to know each component's input and output names to wire them up.
  • It does not choose the right model, database or prompt for you.
  • A pipeline can only answer from documents you have put into the document store.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. repodeepset-ai/haystack, deepset (GitHub) · read 28 Sept 2026
  2. docsIntroduction to Haystack, Haystack documentation (deepset) · read 28 Sept 2026
  3. docsPipelines, Haystack documentation (deepset) · read 28 Sept 2026
  4. docsCreating Pipelines, Haystack documentation (deepset) · read 28 Sept 2026
  5. docsComponents, Haystack documentation (deepset) · read 28 Sept 2026
  6. docsDocument Store, Haystack documentation (deepset) · read 28 Sept 2026
  7. docsAgents, Haystack documentation (deepset) · read 28 Sept 2026