HyDEBuilding with AI

Hypothetical Document Embeddings

4 min readintermediateUpdated 28 Sept 2026
1 · In one line

HyDE has a language model write a made-up answer to your question, then searches for real documents that look like that answer.

1 · What it is

Search by meaning, called dense retrieval, finds documents by comparing embeddings. HyDE, short for Hypothetical Document Embeddings, uses a trick: it searches in a space built only from documents, comparing documents with documents. It was introduced in a paper posted to arXiv in December 2022 and published at the ACL conference in 2023.

HyDE first asks a language model to write a passage that answers the question. The passage is not real and may contain factual errors. That is fine for search. The passage is turned into an embedding, and the search looks for real documents whose embeddings are close to it. The authors expect the encoder to filter most of the wrong details out of the embedding.

Say you search a bike repair manual for why your chain keeps slipping. The model might draft a paragraph blaming a worn chain. That paragraph reads like a page from a repair manual, so the search lands near the real repair pages. Whether the draft guessed the right cause matters less, because the pages you finally read are real ones.

For engineers: the paper used GPT-3 to write the drafts and Contriever to encode them. Contriever is a search model. It learned by comparing pieces of text, without people labelling any examples. HyDE can sample several drafts and average their embeddings. The paper also counts the original question as one of the drafts. It wrote a different instruction for each kind of task. The tests covered finding web pages, answering questions and checking facts, in English and languages such as Swahili, Korean and Japanese. LlamaIndex offers HyDE as a query transform. That step rewrites your question into a hypothetical document before the embedding lookup. Questions that could mean several things are left for future work.

2 · Why it exists

Search by meaning is hard to get right without labelled examples.

Few labelsIt is hard to build a strong meaning-based search system when nobody has marked which documents answer which questions.
New subjectsMany embedding search models do poorly on unfamiliar subject areas.
Weak recallThe search step may miss documents that would have answered the question.
3 · How it works

Follow one question from a fake answer to real documents.

The draft can contain false details, but the documents that come back are real.
  1. 1 · askA question comes in, such as a search typed by a user.
  2. 2 · draftA language model is told to write a passage that answers the question, with no examples given.
  3. 3 · encodeA document encoder turns that passage into an embedding.
  4. 4 · searchThe system finds the stored real documents whose embeddings are most similar.

HyDE searches mainly with an imagined answer, not just the bare question.

4 · Where it's used
WhoWhat they askWhat it works with
Research team“Which papers support or refute this claim?”A drafted paper passage that takes a side
Finance analyst“How are dividends taxed for a small investor?”A drafted financial article passage
New search system“Any question, before there is a search log”A drafted answer passage
5 · What it solves, and what it doesn't
solves
  • It needs no labelled question-and-document pairs, and no model is trained.
  • In its paper it beat the unsupervised Contriever search it was built on.
  • It is worth trying when your documents or questions come from an unusual field.
  • The documents it finds can be passed on to a RAG pipeline's generator.
doesn't solve
  • A short or unclear question can send the draft in the wrong direction.
  • It can push open-ended questions toward one view.
  • In non-English tests it still trailed a search model trained on labelled data.
  • The draft itself is not real and can contain factual errors.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. paperPrecise Zero-Shot Dense Retrieval without Relevance Labels, Gao, Ma, Lin and Callan (arXiv) · read 28 Sept 2026
  2. paperPrecise Zero-Shot Dense Retrieval without Relevance Labels, ACL Anthology · read 28 Sept 2026
  3. repoHyDE: Precise Zero-Shot Dense Retrieval without Relevance Labels, texttron (GitHub) · read 28 Sept 2026
  4. docsHypothetical Document Embeddings (HyDE), deepset (Haystack documentation) · read 28 Sept 2026
  5. docsHyDE Query Transform, LlamaIndex · read 28 Sept 2026
  6. paperUnsupervised Dense Information Retrieval with Contrastive Learning, Izacard et al. (arXiv) · read 28 Sept 2026