Hybrid search
Hybrid search runs a keyword ranker and a vector ranker on one query, then merges the two lists into a single order.
Hybrid search runs two kinds of search on the same query and merges what they find. In Weaviate, the keyword side is BM25. BM25 scores a page by adding up a weight for each query word. Repeating a word in a page helps only up to a limit, which the authors call saturation. The vector side maps the query and each page into one shared space, so it can match meaning when the words differ.
Each side has blind spots. BM25 can only return pages that contain the query’s words. Dense and sparse embedding methods can do much worse than BM25 on tasks they were not trained for. Hybrid search is meant to keep the strengths of both.
The hard part is the merge, because the two scores sit on unrelated scales. Reciprocal rank fusion was published at SIGIR 2009 by Cormack, Clarke and Büttcher. It ignores the scores and uses only positions. Each page gets 1 divided by 60 plus its rank from every list it appears on, and the totals set the final order. The authors fixed 60 during a pilot study. They found the exact value was not critical. Elasticsearch offers this fusion as a built-in way to combine result sets. It also saves you from guessing how to weight the raw scores. Weaviate adds a dial called alpha that sets how much the vector side counts. Alpha runs from 0 to 1. At 0 only keywords count, at 1 only vectors count, and the server uses 0.75 when no value is sent. Weaviate offers rank fusion too, but its default merge works differently. It rescales each side’s raw scores so the best becomes 1 and the worst becomes 0, then adds them. That keeps some of the confidence information that pure rank fusion throws away.
A keyword list and a vector list leave out different pages.
Follow one query that contains a plain phrase and an exact code.
- 1 · matchBM25 sums a weight for each query word that appears in the page.
- 2 · embedA dense ranker places the query and each page in one vector space and orders the pages by closeness.
- 3 · fuseReciprocal rank fusion scores a page by adding 1 over k plus its rank on each list.
- 4 · returnThe caller gets that one order, and a page missing from a list adds nothing from that list.
Fusion mixes ranks. A cross-encoder reranker instead scores the query and a page together after a first search.
| Who | What they ask | What it works with |
|---|---|---|
| Support | “Why does checkout fail with error E-1042?” | Help pages that contain the code or describe a failed payment |
| On-call engineer | “Which runbook covers the timeout after E-1042?” | Notes that print the code or say payment timed out |
| Docs writer | “Where do we explain rotating the API key?” | Guides that name the header or say renew credentials |
| Billing team | “Which article covers a payment that never finished?” | Articles with the error code or a paraphrase of the failure |
- A page that only one ranker returns can still land in the fused order.
- In the original tests, reciprocal rank fusion beat Condorcet, CombMNZ and the best single system by 4% to 5% on average.
- Elasticsearch can fuse a BM25 query with a kNN vector search in one request.
- Weaviate lets you set how much the vector results count against the keyword results.
- A page on neither result list has no rank to add, so fusion cannot bring it in.
- The window size decides how many pages from each list can enter the merge, so a small window can leave good pages out.
- The keyword side still returns a page only when the query's words appear in it.
- Rank fusion throws away how confident each ranker was. Only the position counts.
Sources used
This explainer is written in original language. The links below support its factual claims.
- paperReciprocal Rank Fusion outperforms Condorcet and individual Rank Learning Methods, Cormack, Clarke, and Büttcher, SIGIR 2009 · read 27 Sept 2026
- paperThe Probabilistic Relevance Framework: BM25 and Beyond, Robertson and Zaragoza · read 27 Sept 2026
- docsReciprocal rank fusion, Elasticsearch · read 27 Sept 2026
- docsHybrid search, Weaviate · read 27 Sept 2026
- paperBEIR: A Heterogeneous Benchmark for Zero-shot Evaluation of Information Retrieval Models, Thakur et al. · read 27 Sept 2026