Building with AI

Hybrid search

4 min readintermediateUpdated 28 Sept 2026
1 · In one line

Hybrid search runs a keyword ranker and a vector ranker on one query, then merges the two lists into a single order.

1 · What it is

Hybrid search runs two kinds of search on the same query and merges what they find. In Weaviate, the keyword side is BM25. BM25 scores a page by adding up a weight for each query word. Repeating a word in a page helps only up to a limit, which the authors call saturation. The vector side maps the query and each page into one shared space, so it can match meaning when the words differ.

Each side has blind spots. BM25 can only return pages that contain the query’s words. Dense and sparse embedding methods can do much worse than BM25 on tasks they were not trained for. Hybrid search is meant to keep the strengths of both.

The hard part is the merge, because the two scores sit on unrelated scales. Reciprocal rank fusion was published at SIGIR 2009 by Cormack, Clarke and Büttcher. It ignores the scores and uses only positions. Each page gets 1 divided by 60 plus its rank from every list it appears on, and the totals set the final order. The authors fixed 60 during a pilot study. They found the exact value was not critical. Elasticsearch offers this fusion as a built-in way to combine result sets. It also saves you from guessing how to weight the raw scores. Weaviate adds a dial called alpha that sets how much the vector side counts. Alpha runs from 0 to 1. At 0 only keywords count, at 1 only vectors count, and the server uses 0.75 when no value is sent. Weaviate offers rank fusion too, but its default merge works differently. It rescales each side’s raw scores so the best becomes 1 and the worst becomes 0, then adds them. That keeps some of the confidence information that pure rank fusion throws away.

2 · Why it exists

A keyword list and a vector list leave out different pages.

Missed paraphraseBM25 returns a page only when the query keywords are present, so a paraphrase without those words can stay out.
Meaning, not tokensA dense ranker places words and pages in one shared vector space, so it can match meaning when the exact words differ.
Scores that differThe two rankers give scores on unrelated scales, so their raw numbers cannot be added as they are.
3 · How it works

Follow one query that contains a plain phrase and an exact code.

Exact keyword hits and semantic hits merge into one ranking.
  1. 1 · matchBM25 sums a weight for each query word that appears in the page.
  2. 2 · embedA dense ranker places the query and each page in one vector space and orders the pages by closeness.
  3. 3 · fuseReciprocal rank fusion scores a page by adding 1 over k plus its rank on each list.
  4. 4 · returnThe caller gets that one order, and a page missing from a list adds nothing from that list.

Fusion mixes ranks. A cross-encoder reranker instead scores the query and a page together after a first search.

4 · Where it's used
WhoWhat they askWhat it works with
Support“Why does checkout fail with error E-1042?”Help pages that contain the code or describe a failed payment
On-call engineer“Which runbook covers the timeout after E-1042?”Notes that print the code or say payment timed out
Docs writer“Where do we explain rotating the API key?”Guides that name the header or say renew credentials
Billing team“Which article covers a payment that never finished?”Articles with the error code or a paraphrase of the failure
5 · What it solves, and what it doesn't
solves
  • A page that only one ranker returns can still land in the fused order.
  • In the original tests, reciprocal rank fusion beat Condorcet, CombMNZ and the best single system by 4% to 5% on average.
  • Elasticsearch can fuse a BM25 query with a kNN vector search in one request.
  • Weaviate lets you set how much the vector results count against the keyword results.
doesn't solve
  • A page on neither result list has no rank to add, so fusion cannot bring it in.
  • The window size decides how many pages from each list can enter the merge, so a small window can leave good pages out.
  • The keyword side still returns a page only when the query's words appear in it.
  • Rank fusion throws away how confident each ranker was. Only the position counts.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. paperReciprocal Rank Fusion outperforms Condorcet and individual Rank Learning Methods, Cormack, Clarke, and Büttcher, SIGIR 2009 · read 27 Sept 2026
  2. paperThe Probabilistic Relevance Framework: BM25 and Beyond, Robertson and Zaragoza · read 27 Sept 2026
  3. docsReciprocal rank fusion, Elasticsearch · read 27 Sept 2026
  4. docsHybrid search, Weaviate · read 27 Sept 2026
  5. paperBEIR: A Heterogeneous Benchmark for Zero-shot Evaluation of Information Retrieval Models, Thakur et al. · read 27 Sept 2026