New batches starting this week Β· Limited seats

Hybrid Search and Reranking in RAG: Getting the Right Context Every Time

Most bad RAG answers are lost at retrieval. This practitioner guide shows how to combine BM25 and vector search, fuse results with reciprocal rank fusion, rerank with a cross-encoder and measure every stage.

Hybrid retrieval: keyword BM25 and vector search fused with reciprocal rank fusion, then reranked before reaching the LLM
Last updated Β· 14 min read Β· 3,178 words

Most RAG answers that go wrong were lost at retrieval: the passage holding the answer never reached the model. Hybrid search RAG closes the biggest gap by running keyword search (BM25) and vector search side by side, fusing the two ranked lists with reciprocal rank fusion, and then letting a cross-encoder reranker re-order a short list so the LLM sees the few passages that actually answer the question. This guide covers each stage: how it works, how to tune it, what it costs and how to tell whether it helps.

If you are new to the overall pattern, start with what RAG is and how the pipeline fits together.

Why one retriever is not enough

Vector search and keyword search fail in opposite directions, and enterprise content triggers both failure modes daily.

Where pure vector search misses

Dense vector search compares embeddings, numeric representations of meaning. That is excellent for paraphrase and weak for tokens whose meaning is their exact spelling:

  • Identifiers: policy numbers, ticket IDs, invoice numbers, SKUs. Two IDs that differ by one digit embed almost identically.
  • Codes and numbers: clause 7.3 and clause 7.4 embed almost identically.
  • Acronyms and internal jargon: company abbreviations and project code names were rarely in the embedding model's training data, so their vectors carry little meaning.
  • Names: people, vendors and places, especially transliterated Indian names with several spellings.

Where pure keyword search misses

Keyword search matches words, so it fails when user and document use different vocabulary. "How many days off after my baby is born" may share no important words with a section titled "Parental leave entitlement". Synonyms, paraphrases and conceptual queries favour vector search.

Consider an illustrative case: an IT support assistant for a bank's Hyderabad GCC. Half the questions look like "what does error KB-4471 mean in the settlement batch" and half like "the overnight reconciliation keeps failing after the patch, what should I check". Keyword search nails the first kind, vector search the second. Hybrid retrieval means you stop failing one group of users.

BM25 and keyword search, explained plainly

BM25 is the standard ranking function behind most keyword search engines. It scores a chunk against a query using three ideas:

  1. Term frequency, with saturation. More mentions of "reconciliation" suggest more relevance, but the tenth mention adds far less than the second. The parameter k1 controls how fast the benefit flattens.
  2. Inverse document frequency. Rare words count more. Matching "KB-4471" is strong evidence; matching "process" is weak. This is why BM25 is so good at IDs and codes.
  3. Length normalisation. Long chunks contain more words by chance, so BM25 discounts them relative to the average length, controlled by the parameter b.

Before scoring, text passes through an analyser: tokenisation, lower-casing, often stemming and stop-word removal. A default analyser may split "KB-4471" into two tokens or strip characters from codes, quietly breaking the exact matching you added keyword search for. Test it on your real identifiers. Also note that BM25 scores are unbounded and query-dependent; a score of 14 means nothing on its own, which matters when you fuse.

Hybrid retrieval: run both, then fuse

Send the same query to a keyword index and a vector index in parallel, take the top N from each, and merge them into one ranked candidate list. The hard part is the merge, because the two retrievers score on completely different scales.

Reciprocal rank fusion (RRF)

Reciprocal rank fusion ignores raw scores and uses only positions. Each list a document appears in contributes one divided by a constant plus its rank:

RRF(d) = sum over lists L of 1 / (k + rank_L(d))

rank starts at 1; k is a constant (60 is common)
a document absent from a list gets nothing from it

The constant k dampens the advantage of the very top positions so one retriever's first result cannot dominate. Smaller k rewards top ranks more sharply. The value 60 comes from the original RRF paper and is a sensible default.

Toy worked example (illustrative, k = 60). BM25 returns chunks A, C, D in that order; vector search returns B, C, A.

ChunkBM25 rankVector rankCalculationRRF score
A131/61 + 1/630.03227
C221/62 + 1/620.03226
Bnot returned11/610.01639
D3not returned1/630.01587

Fused order: A, C, B, D. Chunks found by both retrievers rise to the top, which is what you want. A and C are nearly tied, which is why RRF output makes a good candidate list for a reranker rather than a final answer. RRF needs no score calibration and is a few lines of code.

Weighted score fusion and its normalisation pitfalls

The alternative combines actual scores: final = alpha * vector_norm + (1 - alpha) * keyword_norm. It keeps information RRF throws away, such as "the top BM25 hit is far ahead of the rest", but only if the scores are comparable.

  • Min-max per query exaggerates tiny gaps. Toy example: cosine scores of 0.83, 0.82 and 0.81 normalise to 1.0, 0.5 and 0.0, turning a near-tie into a wide spread, while BM25 scores of 18.2, 17.9 and 4.1 become 1.0, 0.98 and 0.0. The fused ranking now reflects normalisation artefacts.
  • Outliers and short lists squash or degenerate the normalised values.
  • Missing documents need a score from the retriever that did not return them; treating that as zero is a choice to make deliberately.
  • Alpha is query-dependent. The best weight for "KB-4471" differs from the best weight for a conceptual question. Some teams route ID-like queries to a keyword-heavy setting.

Start with RRF. Move to weighted fusion only when an evaluation set shows it helps, and tune alpha on that set, not by feel.

Reranking: bi-encoders vs cross-encoders

First-stage retrieval is built for speed over millions of chunks, so it is approximate. Reranking in RAG is a second, careful pass over a short list, typically a few dozen fused candidates, to choose the handful that go into the prompt.

A bi-encoder, which is what your vector search already is, embeds the query and each document separately and compares the vectors. Document vectors are computed once at ingestion and stored in a vector database, so search is fast at scale. But the model never sees query and passage together.

A cross-encoder reranker takes the query and one passage together as a single input, lets the transformer attend across both, and outputs a relevance score. It can notice that a passage mentions the right error code but for a different system, or answers a slightly different question. That joint reading is more accurate and much more expensive: nothing can be precomputed, and you pay one model call per query-passage pair.

AspectBi-encoderCross-encoder
InputQuery and document encoded separatelyQuery and document encoded together
Precompute documents?Yes, at ingestionNo, scored at query time
Practical scaleMillions of chunks via an indexTens to low hundreds per query
Fine distinctionsGoodBetter
Role in RAGFirst-stage retrievalSecond-stage reranking

Hosted rerankers, open models and LLM reranking

Hosted rerank APIs from model and search providers, and rerank models inside cloud AI platforms, are the quickest start: send query and candidates, get scores back. Open cross-encoder models can be self-hosted, which suits data-residency requirements common in Indian banking and healthcare. LLM-based reranking asks a general LLM to score each passage (pointwise) or order a list (listwise). It follows nuanced instructions ("prefer current policy over archived versions") but is slower, costlier and sensitive to passage order. Use it when a dedicated reranker is not good enough, and measure the gain.

Keep the reranker's score. A relevance threshold on the top score lets the system say "I could not find this" instead of answering confidently from weak context. Tune it on labelled data.

Want to build this end to end, with pgvector, full-text search, fusion and a reranker in a real application? Cloudsoft's AI, GenAI and Agentic AI course covers RAG retrieval and evaluation hands-on, in the Ameerpet classroom or live online.

Query rewriting, expansion and multi-query

Users type short, ambiguous queries, often mid-conversation. Improving the query is frequently cheaper than improving the index.

  • Conversational rewrite: turn "and what about for contractors?" into a standalone question using chat history. Without it, follow-ups retrieve nothing useful.
  • Expansion: add synonyms and acronym expansions from a curated glossary. A glossary you control beats an LLM guessing what your acronyms mean.
  • Multi-query: generate a few differently worded versions, retrieve for each, and fuse all lists with RRF. Better recall on vague questions, at the cost of extra retrieval calls.
  • Hypothetical answer (HyDE-style): embed a short generated answer instead of the question, since answers resemble documents. Use it for the vector leg only; generated text can invent IDs.
  • Preserve exact tokens: extract IDs, codes and quoted phrases and pass them through verbatim. A rewrite that "corrects" KB-4471 silently breaks retrieval.

Metadata filters and permission filtering

Metadata filters restrict retrieval to the right slice: document type, region, product, effective date, status = current. Good metadata starts at ingestion; see RAG chunking strategies for attaching section titles, dates and source fields to every chunk.

Permission filtering matters most. Apply access control before or within retrieval, never after generation. Every chunk carries the groups allowed to read it, inherited from the source system, and every retrieval leg filters on the current user's entitlements. If you retrieve unfiltered and clean up later, restricted text has already been ranked, perhaps sent to a third-party reranker, and perhaps written into the answer.

Two details catch teams out. With approximate nearest-neighbour indexes, filtering applied after the vector search can return too few results when the filter is selective; use your engine's filtered search so you still get N permitted candidates. And both legs must apply the same filter: a hybrid system where only the vector leg checks permissions is a data-exposure bug waiting for an audit. More in enterprise AI security.

The full retrieval pipeline

user question + chat history
        |
  query rewrite / expand
  (keep IDs verbatim)
        |
  user entitlements + metadata filter
        |
   +----+-----------------+
   |                      |
 BM25 / full-text     vector search
 top N (filtered)     top N (filtered)
   |                      |
   +----------+-----------+
              |
     fusion (RRF or weighted)
              |
   cross-encoder rerank (top M)
              |
   threshold? -- below --> "not found"
              |
   top K chunks into prompt
              |
     LLM answer + citations

The latency and cost budget

Every stage adds time and money. Set a budget per stage and measure p50 and p95 latency from production traces, not your laptop.

StageWhat drives latency and costMain levers
Query rewriteAn extra LLM call before retrieval startsSmall fast model; skip for standalone queries
Keyword + vectorIndex size, filters, ANN settingsRun in parallel; filtered indexes; sensible N
Multi-queryMultiplies retrieval callsOnly for vague queries; cap variants
RerankNumber and length of candidatesRerank tens, not hundreds; trim long chunks
GenerationPrompt size grows with K and chunk lengthBetter ranking lets you send fewer chunks

A good reranker often saves money overall: it lets you send fewer, better chunks, and generation tokens are usually the most expensive part of a request. See cloud cost optimisation for AI for broader tactics.

How to evaluate each stage

Build a labelled set of real questions mapped to the chunks that answer them, deliberately including ID-heavy queries, paraphrases, acronyms and questions with no answer in the corpus. Then measure stage by stage:

  • Each retriever alone: recall@N. Which leg fails on which query type?
  • After fusion: recall@M and MRR. Fusion should not lose chunks either leg found.
  • After reranking: precision and nDCG at K, and how often the right chunk is first.
  • Threshold: correct declines on no-answer questions versus wrong declines on answerable ones.
  • End to end: context precision, context recall, faithfulness and answer relevance, for example with Ragas.

Run every change (new analyser, different k, new reranker, more candidates) as an ablation against the same set. The RAG evaluation metrics guide explains each metric and how to build the test set.

PostgreSQL full-text search plus pgvector. A strong default if your application already runs on PostgreSQL. Full-text search uses tsvector columns with GIN indexes; pgvector stores embeddings with HNSW or IVFFlat indexes. Run both queries with the same WHERE clause for permissions and metadata, then fuse with RRF in SQL or application code. PostgreSQL's built-in ranking is not BM25, though extensions can add BM25-style ranking. The enterprise knowledge assistant project uses this approach.

OpenSearch hybrid. Mature BM25 search, analysers and filtering alongside k-NN vector search. A hybrid query runs both and a search pipeline normalises and combines scores, with rank-based fusion available in recent versions.

Azure AI Search. A managed service supporting keyword, vector and hybrid queries, merging hybrid results with RRF, plus an optional semantic ranker that reranks top results. Filter-based security trimming fits Microsoft Entra ID groups, making it natural for teams on Azure OpenAI.

Common mistakes

  • Adding a reranker before fixing chunking. It cannot rescue an answer split across two chunks.
  • Analysers that mangle IDs, defeating the reason you added keyword search.
  • Fusing raw scores. Adding BM25 to cosine without calibration ranks by scale, not relevance.
  • Permission filter on one leg only, or applied after generation.
  • No threshold, so weak retrieval becomes confident wrong answers.
  • Ignoring stale versions: without status and date filters, last year's policy ranks happily.
  • Tuning without a test set, so you cannot tell improvement from noise.

Taking a pipeline like this from a notebook into a customer's environment, with their identity system, documents and acceptance criteria, is a core part of what Forward Deployed Engineers do.

FAQ

What is hybrid search in RAG?

Hybrid search in RAG runs keyword search, such as BM25 or database full-text search, and vector search on the same query, then merges the two ranked lists. Keyword search reliably finds exact terms like IDs, error codes and names, while vector search finds paraphrases, so together they retrieve the right context more often than either alone.

Neither is better in general. BM25 is stronger for exact terms, rare words, identifiers and numbers. Vector search is stronger when the user and the document use different words for the same idea. Enterprise content attracts both kinds of query, which is why most production RAG systems combine them.

What is reciprocal rank fusion?

Reciprocal rank fusion merges ranked lists using positions instead of scores. Each document receives 1 / (k + rank) from every list it appears in, and the contributions are added; k is commonly 60. Documents ranked well by both retrievers rise to the top, and no score normalisation is needed.

What is the difference between a bi-encoder and a cross-encoder?

A bi-encoder embeds the query and each document separately, so document vectors can be precomputed and searched quickly at scale. A cross-encoder reads the query and a document together and outputs a relevance score, which is more accurate but must run at query time for every pair. Bi-encoders handle first-stage retrieval; cross-encoders rerank a short list.

Do I always need a reranker in RAG?

Not always, but it is often one of the cheapest quality gains once chunking and hybrid retrieval are sound. Add one when evaluation shows the correct chunk is usually in the top few dozen candidates but not the top few. If the correct chunk is not retrieved at all, fix chunking, metadata or retrieval first.

Where should permission filtering happen in a RAG pipeline?

Before or within retrieval, on every retrieval leg. Each chunk should carry the groups allowed to read it, and both the keyword and vector queries should filter on the current user's entitlements. Filtering after generation is too late, because restricted text may already have been ranked, sent to a reranker or used in the answer.

Can I do hybrid search with PostgreSQL?

Yes. PostgreSQL full-text search handles the keyword side and the pgvector extension handles vector search. Run both with the same permission and metadata filters, then fuse with reciprocal rank fusion. Built-in PostgreSQL ranking is not BM25, but extensions can add BM25-style ranking if you need it.

Build production-grade retrieval

Hybrid retrieval, reranking, permission-aware filtering and stage-by-stage evaluation separate a RAG demo from an assistant people trust. To practise them on real cloud AI platforms with guided labs, explore Cloudsoft's AI, GenAI and Agentic AI training in Hyderabad, in our Ameerpet classroom or live online. Call +91 96660 19191 for a free demo.

Share𝕏infβœ‰
EnrollWhatsAppCall us