Reranking definition
Reranking is a second retrieval stage that takes the top candidates from an initial search and reorders them with a more accurate but slower model, usually a cross-encoder that reads the query and each document together. It improves precision at the top of the results, which matters most in RAG systems that pass only a few documents to an LLM.
Why search needs a second stage
First-stage retrieval must search millions of documents in milliseconds, so it uses fast methods: keyword indexes and vector similarity between precomputed embeddings. Those methods are good at finding a relevant set but imperfect at ordering it. The single best answer might sit at position 12, below documents that merely share vocabulary with the query.
Reranking fixes the order. It takes perhaps the top 50 or 100 candidates, scores each with a model that examines the query and document together, and returns the best few. Because it runs on a small set, it can afford a heavier model without slowing the overall search much.
In a RAG system this matters a great deal: the language model usually sees only the top five to ten chunks, so whatever ranks at position 11 might as well not exist. Better ordering means the answer is grounded in the right evidence.
Bi-encoders vs cross-encoders
Embedding models used for vector search are bi-encoders: they encode the query and each document separately into vectors, which is why document vectors can be computed in advance. That speed has a cost, since the model never sees the query and document side by side and cannot weigh how specific words relate to each other.
Cross-encoders take the query and a document as one input and output a relevance score, letting the model attend to interactions between them. They judge relevance much more accurately but are far too slow to run against a whole corpus, which is exactly why they rerank a shortlist. LLMs can also act as rerankers, at higher cost.
Reranking options
Teams can choose between hosted APIs, open models they run themselves and built-in features of search platforms. Common options include the following, and most can be swapped with little code change if evaluation favors another:
- Hosted rerank APIs such as Cohere Rerank, Voyage AI and Jina AI rerankers
- Open cross-encoder models, such as BGE rerankers and MiniLM models trained on MS MARCO, served with sentence-transformers
- Built-in semantic ranking in platforms such as Azure AI Search, Elasticsearch and Vertex AI Search
- LLM-based reranking, where a language model scores or orders candidates for complex relevance judgments
Using reranking well
Rerank a candidate set large enough to contain the right answers, often 25 to 100 items from hybrid search, then keep the top few. Measure with ranking metrics such as recall at k and NDCG on a labeled set of real queries, and watch latency, since reranking typically adds tens to hundreds of milliseconds depending on the model and candidate count.
Reranking cannot rescue poor first-stage retrieval: if the right document is not in the candidate set, no reranker will find it. Nexzem tunes both stages together in RAG development projects, adding a reranker when evaluation shows it improves answer quality enough to justify the extra latency and cost.