Query Optimization for RAG: Rewriting, Expanding, Resolving, and Decomposing User Queries
This article details four query optimization strategies—rewrite, expansion, resolution, and decomposition—with code examples and benchmarks showing how each step improves retrieval recall from 62% to 87% and accuracy from 55% to 88% in a RAG pipeline.
Why Query Optimization Matters
Users rarely phrase questions using the exact terminology found in documentation. They use colloquial language, make typos, and rely on pronouns like "it" or "that" in multi-turn conversations. A raw query such as "How do I use that feature mentioned earlier?" contains no specific technical terms, causing both vector and BM25 retrieval to fail. After optimization, the same intent becomes multiple precise queries: "XX feature usage", "XX feature configuration", "XX feature steps", which successfully hit relevant documents.
Four Directions of Query Optimization
Rewrite : Convert colloquial to formal terminology, correct typos, expand abbreviations.
Expansion : Generate synonym variants, related concepts, and multi-angle questions.
Resolution : Resolve coreferences ("it", "this"), complete ellipses, merge multi-turn context.
Decomposition : Break complex questions into 2-4 single-focus sub-questions for multi-hop reasoning.
Query Rewriting
Colloquial to Formal
A Spring service uses an LLM with a system prompt that enforces five rules: preserve original intent, convert colloquialisms to technical terms, fix typos, expand abbreviations, and output only the rewritten query.
@Service
public class QueryRewriter {
private final ChatClient chatClient;
public QueryRewriter(ChatClient.Builder builder) {
this.chatClient = builder
.defaultSystem("""
You are a query rewriting assistant. Rewrite user colloquial questions into standard technical queries.
Rules:
1. Preserve original meaning, do not add non-existent information
2. Convert colloquial expressions to technical terminology
3. Correct typos
4. Expand abbreviations
5. Output only the rewritten query, no explanation
""")
.build();
}
public String rewrite(String originalQuery) {
return chatClient.prompt(originalQuery).call().content();
}
}Examples:
"How to do that login" → "Login feature usage method"
"Why can't I find the order" → "Order query failure causes and solutions"
"What does this error mean" → "Error message meaning and solution"
Typo Correction
A dedicated method sends the query to the LLM with a prompt to correct typos while preserving meaning, returning only the corrected query.
public String correctTypos(String query) {
String prompt = """
Correct typos in the following query, keep original meaning unchanged:
Query: %s
Only output the corrected query.
""".formatted(query);
return chatClient.prompt(prompt).call().content();
}Query Expansion
Synonym Expansion
The QueryExpander service prompts the LLM to generate three semantically equivalent but differently phrased versions to increase recall.
@Service
public class QueryExpander {
private final ChatClient chatClient;
public List<String> expand(String query) {
String prompt = """
Generate 3 versions with same semantics but different expressions for the following query to improve retrieval recall.
Original query: %s
Requirements:
1. Each version uses different phrasing
2. Include synonyms and near-synonyms
3. One per line, no numbering
Output:
""".formatted(query);
String result = chatClient.prompt(prompt).call().content();
return Arrays.asList(result.split("
"));
}
}Example: "How to configure database connection" expands to "Database connection configuration method", "How to set up database connection pool", "Data source configuration steps".
HyDE (Hypothetical Document Embeddings)
HyDE first generates a hypothetical answer, then uses that answer for retrieval. Since questions are short while documents are long, the hypothetical answer's phrasing aligns better with document style, improving vector similarity.
@Service
public class HydeQueryOptimizer {
private final ChatClient chatClient;
private final VectorStore vectorStore;
public List<Document> hydeSearch(String query) {
// 1. Generate hypothetical answer
String hypotheticalAnswer = chatClient.prompt("""
Based on the following question, generate a possible answer (even if uncertain):
Question: %s
Generate a brief, possible answer:
""".formatted(query)).call().content();
// 2. Retrieve using hypothetical answer
return vectorStore.similaritySearch(
SearchRequest.builder()
.query(hypotheticalAnswer)
.topK(10)
.build()
);
}
}Query Resolution
Coreference Resolution
In multi-turn dialogues, pronouns like "it", "this", "that" refer to earlier entities. The CoreferenceResolver feeds conversation history and the current query to an LLM, instructing it to replace pronouns with concrete referents and complete omitted subjects.
@Service
public class CoreferenceResolver {
private final ChatClient chatClient;
public String resolve(String currentQuery, List<Message> history) {
String historyText = history.stream()
.map(m -> m.getMessageType() + ": " + m.getText())
.collect(Collectors.joining("
"));
String prompt = """
Based on conversation history, replace pronouns in current question with specific content.
Conversation history:
%s
Current question: %s
Requirements:
1. Replace "it", "this", "that" with concrete content
2. Complete omitted subjects
3. Output only the rewritten complete question
""".formatted(historyText, currentQuery);
return chatClient.prompt(prompt).call().content();
}
}Examples:
History: "What is Product A's warranty policy?" → Current: "How often does it need maintenance?" → Resolved: "How often does Product A need maintenance?"
History: "How to configure database?" → Current: "How to change the port?" → Resolved: "How to change the port in database configuration?"
Multi-Query Retrieval
Core Idea
Retrieve with multiple rewritten queries, then merge results using Reciprocal Rank Fusion (RRF). The MultiQueryRetriever orchestrates rewriting, expansion, and hybrid search.
@Service
public class MultiQueryRetriever {
@Autowired private QueryRewriter rewriter;
@Autowired private QueryExpander expander;
@Autowired private HybridRetriever hybridRetriever;
public List<Document> multiQuerySearch(String originalQuery, int topK) {
// 1. Rewrite original query
String rewritten = rewriter.rewrite(originalQuery);
// 2. Expand to multiple queries
List<String> queries = new ArrayList<>();
queries.add(originalQuery);
queries.add(rewritten);
queries.addAll(expander.expand(rewritten));
// 3. Retrieve per query
Map<String, Double> scores = new HashMap<>();
Map<String, Document> docMap = new HashMap<>();
for (String query : queries) {
List<Document> results = hybridRetriever.hybridSearch(query, topK);
for (int i = 0; i < results.size(); i++) {
Document doc = results.get(i);
double score = 1.0 / (60 + i + 1);
scores.merge(doc.getId(), score, Double::sum);
docMap.putIfAbsent(doc.getId(), doc);
}
}
// 4. Sort by fused score
return scores.entrySet().stream()
.sorted(Map.Entry.<String, Double>comparingByValue().reversed())
.limit(topK)
.map(e -> docMap.get(e.getKey()))
.collect(Collectors.toList());
}
}RAG-Fusion
RAG-Fusion generates multiple queries from different angles, retrieves in parallel, and fuses via RRF.
public List<Document> ragFusion(String originalQuery, int topK) {
// 1. Generate multiple queries
List<String> queries = generateQueries(originalQuery);
// 2. Retrieve per query in parallel
List<List<Document>> allResults = queries.parallelStream()
.map(q -> hybridRetriever.hybridSearch(q, topK))
.collect(Collectors.toList());
// 3. RRF fusion
return rrfFusion(allResults, topK);
}
private List<String> generateQueries(String query) {
String prompt = """
Generate 4 queries from different angles for comprehensive retrieval.
Original question: %s
Requirements:
1. Ask from different angles
2. Use different keywords
3. One per line
Output:
""".formatted(query);
String result = chatClient.prompt(prompt).call().content();
return Arrays.asList(result.split("
"));
}Query Decomposition
Complex questions are split into 2-4 simple sub-questions, each targeting a single information point without overlap.
@Service
public class QueryDecomposer {
private final ChatClient chatClient;
public List<String> decompose(String complexQuery) {
String prompt = """
Decompose the following complex question into 2-4 simple sub-questions for separate retrieval.
Complex question: %s
Requirements:
1. Each sub-question covers one information point
2. Sub-questions must not overlap
3. One per line
Output:
""".formatted(complexQuery);
String result = chatClient.prompt(prompt).call().content();
return Arrays.asList(result.split("
"));
}
}Example: "What are the differences between Product A and Product B in warranty, price, and after-sales?" decomposes to: 1. Product A warranty and price 2. Product B warranty and price 3. Product A vs Product B after-sales comparison
Complete Query Optimization Pipeline
The QueryOptimizationPipeline chains all steps: coreference resolution → rewrite → expand → multi-query retrieval → rerank.
@Service
public class QueryOptimizationPipeline {
@Autowired private CoreferenceResolver coreferenceResolver;
@Autowired private QueryRewriter rewriter;
@Autowired private QueryExpander expander;
@Autowired private MultiQueryRetriever multiQueryRetriever;
@Autowired private Reranker reranker;
public List<Document> optimizeAndRetrieve(String query, List<Message> history) {
// 1. Coreference resolution
String resolved = history.isEmpty()
? query
: coreferenceResolver.resolve(query, history);
// 2. Rewrite
String rewritten = rewriter.rewrite(resolved);
// 3. Expand
List<String> expanded = expander.expand(rewritten);
// 4. Multi-query retrieval
List<Document> candidates = multiQueryRetriever.multiQuerySearch(rewritten, 20);
// 5. Rerank
return reranker.rerank(rewritten, candidates, 5)
.stream()
.map(RerankResult::getDocument)
.collect(Collectors.toList());
}
}Effect Comparison
Incremental improvements measured on recall and accuracy:
No optimization: Recall 62%, Accuracy 55%
+ Rewrite: Recall 72%, Accuracy 65%
+ Expansion: Recall 80%, Accuracy 71%
+ Multi-query retrieval: Recall 87%, Accuracy 78%
+ Rerank: Recall 87%, Accuracy 88%
Summary: Each step adds value; reranking provides the largest accuracy boost.
Next Episode Preview
Episode 9 will cover Prompt Engineering for RAG: core prompt structure, constraining hallucinations, citing sources, and handling "not found" cases.
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
Coder Trainee
Experienced in Java and Python, we share and learn together. For submissions or collaborations, DM us.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
