Building with AI

Reranking

4 min readintermediateUpdated 28 Sept 2026
1 · In one line

Reranking takes the short list a first search returned and scores each item again with a slower, more careful model.

1 · What it is

Reranking is the second pass of a search. The first pass uses a cheap method such as BM25 to gather many possibly relevant documents. The second pass, reranking, scores each of those documents with a heavier model and reorders them. Reranking is one of the last steps before results are shown or passed on.

Rerankers come in two main types. A cross-encoder reads the question and one passage as a single joined input. A bi-encoder instead makes one vector for the question and one for each document, then compares them with a dot product or cosine. In the 2019 BERT reranker, the question went in as the first sentence and the passage as the second. The model then output the probability that the passage was relevant.

Cohere sells reranking as a hosted model that sorts the texts you send by how relevant they are to your query. Its current list starts with rerank-v4.0-pro and a lighter rerank-v4.0-fast. The older rerank-v3.5 has a window of 4,096 tokens for each document, and the query uses up part of that window too. When a query plus one document runs past it, Cohere does not reject the pair. It splits the document into pieces and scores them in several passes.

The price is speed. In the BEIR benchmark, dense retrievers ran 20 to 30 times faster than rerankers. The reranker still beat BM25 on 16 of the 18 datasets.

2 · Why it exists

A fast first search gathers candidates cheaply, and a slower model can then judge each one more carefully.

Rough first listA keyword method such as BM25 can pull in a big pile of maybe-relevant documents, around a thousand in one classic setup.
Vectors stay apartA bi-encoder turns the question and each document into vectors on their own, so neither vector knows about the other.
Misses stay outSome relevant passages never make the first list at all.
3 · How it works

Follow one password question from a rough list to a new order.

The highlighted step reads the question and one passage as a single input. The right-hand list is the order after those pair scores.
  1. 1 · retrieveA fast first search gathers the candidates, for example with BM25.
  2. 2 · scoreA cross-encoder reads the question and one passage joined together as a single input and outputs a relevance score.
  3. 3 · sortEvery passage gets its own score, and the list is sorted by those scores.

The key move is one input: the question and the passage are read together.

4 · Where it's used
WhoWhat they askWhat it works with
Support lead“Which help article answers this refund question?”The top hits from the help centre
Lab researcher“Which paper matches this methods question?”The first-stage list of papers
IT help“How do I reset a forgotten password?”Handbook passages from keyword search
Shopper“Which item fits this short product question?”Candidate product descriptions
5 · What it solves, and what it doesn't
solves
  • It can move the best passage to the top of a list the first search already found.
  • Reading the question and passage together lets the model judge how well that exact pair fits.
  • It can sit after keyword search, vector search or a mix of the two.
  • On the MS MARCO passage task, a BERT reranker beat the previous best system by 27% on the MRR@10 measure.
doesn't solve
  • It cannot find what the first search missed. Azure's semantic ranker, for example, does not search the whole collection again.
  • It works on a cutoff. Azure's semantic ranker only reorders the top 50 results from the first ranking.
  • It is slow. In the BEIR tests, reranking the top 100 BM25 hits was among the slowest options.
  • It can still lose on unfamiliar tasks. A cross-encoder trained on MS MARCO fell behind BM25 on two very different datasets.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. paperPassage Re-ranking with BERT, Nogueira and Cho · read 27 Sept 2026
  2. docsCohere's Rerank Model (Details and Application), Cohere · read 27 Sept 2026
  3. docsSemantic reranking, Elasticsearch · read 27 Sept 2026
  4. docsSemantic Ranking Overview, Microsoft · read 27 Sept 2026
  5. paperBEIR: A Heterogeneous Benchmark for Zero-shot Evaluation of Information Retrieval Models, Thakur et al. · read 27 Sept 2026