RAG Retrieval Quality: Hybrid Search, RRF Fusion, Rerank & Query Rewriting

This article explains how to improve RAG retrieval quality by combining vector and keyword search via RRF fusion, adding cross-encoder reranking, rewriting user queries for better retrieval, and using HyDE to generate hypothetical answers for embedding-based search, with a practical Java implementation example.

Dabaoshi
Dabaoshi
Dabaoshi
RAG Retrieval Quality: Hybrid Search, RRF Fusion, Rerank & Query Rewriting

1. Where Pure Vector Search Falls Short

Vector embeddings capture statistical semantic similarity, but they struggle with three categories of queries:

Exact-match information : order IDs, product SKUs, model numbers, proper nouns. These strings lack semantic meaning; vector models cannot pinpoint a specific character sequence the way keyword search does.

Negation and opposite semantics : "How to request a refund" and "How to reject a refund request" sit close in vector space because both discuss refunds, yet they require opposite answers.

Long-tail rare terms : a term appearing only once in a single document may have an imprecise embedding if the training data rarely contains it.

These weaknesses are exactly where traditional keyword search (e.g., BM25) excels — BM25 matches literal token overlap, so a precise order ID or unique term is reliably retrieved if present in the document. The core idea is not to replace keyword search with vector search, but to combine them so each covers the other's blind spots.

2. Hybrid Search: Vector and Keyword Each Play Their Role

Hybrid search runs the same query through both a vector retriever and a keyword retriever (such as BM25), producing two ranked candidate lists, then merges them into a single ranking.

User query: "When will order 202408290012 arrive?"
         │
         ├──→ Vector search: embed query, find semantically related docs
         │     Candidates: ["Logistics delivery time policy", "Common logistics FAQ", ...]
         │
         ├──→ Keyword search (BM25): match literal token overlap
         │     Candidates: ["Order 202408290012 processing record", "Logistics delivery time policy", ...]
         │
         ▼
   Result fusion: merge, deduplicate, re-rank the two candidate sets

The two retrievers may overlap (e.g., "Logistics delivery time policy" appears in both, signaling both semantic and lexical relevance) and each may have unique hits (keyword search nails the exact order record; vector search finds semantically relevant policy docs the keyword search misses). Hybrid search's value lies in this complementarity: semantic search covers "same meaning, different wording"; keyword search covers "must match exactly".

3. Result Fusion: How RRF Merges Two Rankings

The two retrievers produce scores on completely different scales — vector search yields cosine similarity (0.1–1.0), while BM25 scores vary with document length and term frequency and have no cross-query normalization. Simply adding weighted scores is unsound because they are not on the same scale.

A more robust approach uses only rank , not raw scores — this is RRF (Reciprocal Rank Fusion) . Its formula is simple: RRF_score(doc) = Σ 1 / (k + rank_i) For each retriever, a document at position rank (1-based) contributes 1 / (k + rank) to its fused score. Summing contributions across all retrievers gives the final RRF score. The smoothing constant k is commonly set to 60 — an empirical value from early IR research that has proven stable across many datasets and is the default in several mainstream search engines and databases.

Doc A: vector rank 1, keyword rank 3
   RRF(A) = 1/(60+1) + 1/(60+3) = 0.0164 + 0.0159 = 0.0323
Doc B: vector miss (not in candidates), keyword rank 1
   RRF(B) = 0 + 1/(60+1) = 0.0164

RRF's advantage is that it ignores score magnitudes entirely, relying only on ranks, making implementation simple and robust — hence its adoption as the default fusion strategy in many off-the-shelf retrieval systems.

Version note : k=60 is a widely used empirical default, not the only theoretically correct value. Smaller k amplifies differences between top ranks; larger k makes fusion approach simple rank addition. The choice of fusion algorithm and k varies by database/search engine; verify the target product's current documentation when integrating.

4. Rerank: Why Fuse Then Re-Rank Again

After fusion, results are much better than single-retriever output, but both vector search and BM25 share a limitation: they score without ever jointly examining the query and the full document semantics. They are Bi-Encoders — query and document are encoded independently, then similarity is computed. This allows pre-computing document vectors for fast retrieval.

Rerank introduces a Cross-Encoder : the query and each candidate document are concatenated and fed together into the model, which directly judges relevance. Cross-Encoder judgments are typically more accurate than Bi-Encoder similarity scores, but they cannot pre-compute document representations — each query-candidate pair must be scored at query time, so reranking is only feasible on a small candidate set (e.g., top 20–50), not the entire corpus.

Stage 1 (Recall): Vector + Keyword search, each take Top 20~50, optimize for recall (don't miss)
                  │
                  ▼
Stage 2 (Rerank): Cross-Encoder scores each candidate finely, re-sorts
                  │
                  ▼
            Take Top 5, feed into generation context

This is the classic funnel design: stage 1 maximizes recall (better to over-retrieve than miss), stage 2 maximizes precision on the small candidate pool. Because reranking only processes dozens of candidates, even a heavier model keeps overall latency manageable.

5. Query Rewriting: User Phrasing ≠ Retrieval-Friendly Phrasing

Previous sections optimize the retrieval action itself; another class of problems occurs before retrieval — the user's raw query is often not a good search query.

Raw user input: "This thing doesn't work anymore, so frustrating"

High emotion, vague references ("this thing", "doesn't work") — direct retrieval will likely fail. Query Rewriting uses a model to transform such input into a clearer, retrieval-friendly form:

Rewritten: "After-sales process for product malfunction"

Common rewriting scenarios go beyond removing emotional language:

Resolve references : in multi-turn dialogue, "this" or "it" must be resolved to a concrete product name or order ID using conversation history.

Decompose compound questions : a single sentence may ask two things ("Can I return this, and who pays shipping?") — split into two independent searches, retrieve answers separately, then merge.

Expand synonymous expressions : generate several paraphrases of the original query, search each, and union the results to increase recall coverage.

6. HyDE: Let the Model "Imagine" an Answer First

HyDE (Hypothetical Document Embeddings) is a counter-intuitive but effective technique: instead of searching with the user's question, first ask the model to "pretend" it knows the answer and generate a hypothetical answer, then use that hypothetical answer's embedding for vector search.

User question: "How long is the warranty for this keyboard?"
         │
         ▼
Model generates a hypothetical answer (need not be factual, just "answer-like"):
"This product comes with a one-year warranty; non-user damage is repaired free within the warranty period..."
         │
         ▼
Search the knowledge base with the hypothetical answer's vector, not the raw question's vector

Why does this work? Because user questions and knowledge-base documents often differ in linguistic style — questions are short and colloquial; documents are formal and declarative. Vector search compares semantic direction; a hypothetical "answer-style" text is stylistically closer to real documents, so its embedding often retrieves more relevant documents than the raw question's embedding. HyDE adds an extra generation step (latency and cost), so it suits scenarios where the style gap between queries and documents is large.

7. Hands-On: Adding Hybrid Search to a Product Knowledge Base

Hybrid search architecture diagram
Hybrid search architecture diagram

Assemble the core ideas from sections 3–4 — vector search + keyword search + RRF fusion + rerank — into a replacement retrieve method for the vector-only version in article 5:

public List<DocChunk> retrieveHybrid(long productId, String question, int topK) {
    List<ScoredChunk> vectorResults = vectorSearch(productId, question, 30); // Vector top 30
    List<ScoredChunk> keywordResults = bm25Search(productId, question, 30);   // Keyword top 30
    Map<String, Double> rrfScores = new HashMap<>();
    fuseByRank(vectorResults, rrfScores);   // Accumulate 1/(60+rank) per rank
    fuseByRank(keywordResults, rrfScores);
    List<DocChunk> fused = rrfScores.entrySet().stream()
        .sorted(Map.Entry.<String, Double>comparingByValue().reversed())
        .limit(20)                          // Fused top 20 enter rerank
        .map(e -> chunkById(e.getKey()))
        .toList();
    return rerankClient.rerank(question, fused, topK); // Cross-Encoder rerank, final topK
}

private void fuseByRank(List<ScoredChunk> results, Map<String, Double> rrfScores) {
    int k = 60;
    for (int rank = 0; rank < results.size(); rank++) {
        String id = results.get(rank).chunkId();
        rrfScores.merge(id, 1.0 / (k + rank + 1), Double::sum);
    }
}

This implementation mirrors the funnel structure from sections 3–4: vector + keyword each take 30 → RRF fusion takes top 20 → rerank produces final top K. Whether to add query rewriting and HyDE depends on whether observed retrieval failures match those specific problems — which is exactly why the next article's evaluation framework matters: without failure-case analysis, optimization becomes "adding features by gut feeling".

8. How to Combine These Techniques Depends on the Scenario

This article covers four technique families — hybrid search, RRF fusion, rerank, query rewriting/HyDE — but not every RAG system needs all of them . Selection depends on cost budget and scenario characteristics:

If the knowledge base contains many exact IDs, model numbers (order IDs, SKUs) , hybrid search usually yields clear gains.

If candidate quality is uneven and mis-retrieval cost is high (e.g., customer service where wrong answers cause complaints), rerank often has high ROI.

If user inputs are typically short, colloquial, and ambiguous , investing in query rewriting pays off.

If the goal is simplest viable implementation, data volume is small, and mis-retrieval is tolerable , the pure vector search from article 5 is perfectly adequate — no need to adopt the full stack from day one.

Let cost and scenario drive decisions; do not treat this article as a mandatory production checklist.

Three takeaways:

Vector search and keyword search are complementary, not substitutes; exact identifiers almost always need keyword search as a backstop.

RRF uses only ranks, ignoring score scales, making it the simplest reliable fusion method; k=60 is a widely validated empirical default, not the only answer.

Rerank applies a more precise but expensive Cross-Encoder to a small candidate set — the classic "recall first, precision second" funnel design.

Next up: AI Application 07 · RAG Effectiveness Evaluation: How to Know If Your QA Bot Is Reliable . After all these optimizations, how do we quantify whether things actually improved, rather than just feeling better? The next article provides measurable, regression-testable evaluation methods.

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

RAGBM25Query RewritingRerankHybrid SearchHyDERRFCross-Encoder
Dabaoshi
Written by

Dabaoshi

Practical utilities

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.