Semantic search
Semantic search finds text by meaning, so a passage can match a question even when the two share no words.
Semantic search finds text by what it means, not only by the exact words it uses. The question and every document are turned into vectors that can be compared directly. Often one embedding model makes both. DPR instead uses two separate encoders, trained together so that the dot product of their outputs ranks passages well. Sentences with similar meanings end up close together in one vector space. The search returns the documents whose vectors sit nearest the question’s vector.
This matters when people and documents use different words for the same thing. The Dense Passage Retrieval paper, DPR for short, shows a film question that asks for the “bad guy”, while the passage holding the answer only says “villain”. A word-matching index would have trouble linking those two texts. In dense retrieval, texts built from completely different words can still land near each other. DPR scores closeness with the dot product of the question vector and the passage vector.
For large collections, OpenAI suggests keeping the vectors in a vector database. OpenAI’s embeddings are normalized to length 1. For those vectors, a plain dot product gives cosine similarity a little faster. Elasticsearch ranks documents by how similar their vector field is to the query vector. A larger similarity score puts a document higher.
A keyword index can miss the right passage when the question and the text use different words.
Follow one question from the wording to the nearest passage.
- 1 · embedOne encoder turns each passage into a vector, and a second encoder does the same for the question.
- 2 · storeA large set of those vectors is searched in a vector database.
- 3 · scoreSimilarity is the dot product of the question vector and the passage vector.
- 4 · returnThe passages that land nearest the question come back first.
When the vectors have length 1, cosine similarity is the same comparison as that dot product.
| Who | What they ask | What it works with |
|---|---|---|
| Film student | “Who is the bad guy in that trilogy?” | A biography that says the actor played the villain |
| Shopper | “What cord charges this old laptop?” | A manual line that says USB power adapter |
| Help desk | “Why does the nightly job stop?” | A runbook that says the batch exits early |
| Librarian | “Do you have a page on the scoundrel in that fantasy film?” | A chapter that never uses the word scoundrel |
- A passage with no shared words can still be the right hit.
- Sentences with a similar meaning sit close in one vector space. Their embeddings can be compared with cosine similarity.
- A search can return results based on meaning, not just on keywords.
- Passage vectors can be kept in an index built for fast similarity search.
- When a question copies the passage's own words, keyword ranking has a clear edge.
- Rare exact names are a weak spot. In one DPR test, the name Thoros of Myr was the key, and the dense retriever missed it.
- The question and the documents must share one vector space. That means the same embedding model, or a question encoder trained together with the document encoder.
- Close does not mean correct. A small distance says two texts are related, not that one answers the other.
Sources used
This explainer is written in original language. The links below support its factual claims.
- paperDense Passage Retrieval for Open-Domain Question Answering, Karpukhin et al. · read 27 Sept 2026
- paperSentence-BERT: Sentence Embeddings using Siamese BERT-Networks, Reimers and Gurevych · read 27 Sept 2026
- docsVector embeddings, OpenAI · read 27 Sept 2026
- docsSemantic search for text, Elastic · read 27 Sept 2026
- docsDense vector field type, Elastic · read 27 Sept 2026