RAG Prompt Engineering: Constraining Models to Use Retrieved Knowledge Correctly
This article details how to design effective RAG prompts that prevent hallucination, enforce source citation, handle missing information, concatenate multiple documents with metadata, manage multi-turn conversations with history compression, and provides a complete prompt template with a tuning checklist.
Core Structure of RAG Prompts
Bad vs Good Prompt Examples
A minimal prompt like
String badPrompt = "根据以下内容回答问题:
" + context + "
问题:" + question;lacks constraints, format requirements, and handling for missing information.
A robust prompt defines role, context, question, and explicit requirements:
String goodPrompt = """
你是一个基于知识库的问答助手。请严格根据以下资料回答用户问题。
## 资料
%s
## 用户问题
%s
## 回答要求
1. 只使用资料中的信息回答,不要添加资料外的内容
2. 如果资料中没有相关信息,明确回答"根据现有资料无法回答该问题"
3. 回答要简洁、准确、有条理
4. 如果引用了具体内容,标注来源编号
## 输出格式
答案:...
来源:[1][2]
""".formatted(context, question);Constraining Models to Prevent Hallucination
Three Constraint Levels
Strict : "Only use information from the provided materials, do not add any external content" — applicable to legal, medical, financial domains.
Lenient : "Prioritize information from materials, may supplement with common knowledge appropriately" — applicable to general Q&A.
Open : "Refer to the following materials to answer the question" — applicable to creative scenarios.
Handling "Not Found" Cases
The most common RAG pitfall: when materials lack the answer, the model fabricates. Solution: treat "not found" as a normal output, not a failure.
private static final String STRICT_PROMPT = """
你是一个严格基于知识库的问答助手。
## 资料
{context}
## 用户问题
{question}
## 核心规则
1. 只使用资料中的信息回答
2. 如果资料中没有答案,必须回答"根据现有资料,我无法回答这个问题"
3. 禁止编造任何资料中没有的信息
4. 禁止使用你的预训练知识补充回答
## 回答
""";Post-Generation Hallucination Detection
A secondary check using another LLM call:
@Service
public class HallucinationDetector {
private final ChatClient chatClient;
public boolean isHallucination(String answer, String context) {
String prompt = """
判断以下回答是否完全基于提供的资料。
资料:%s
回答:%s
如果回答中包含了资料中没有的信息,输出"是";
如果回答完全基于资料,输出"否"。
只输出"是"或"否"。
""".formatted(context, answer);
String result = chatClient.prompt(prompt).call().content();
return result.contains("是");
}
}Citing Sources
Why Cite Sources
Verifiable : Users can verify answer correctness.
Credibility : Sourced answers are more trustworthy.
Traceable : Issues can be traced to original documents.
User Experience : Users can jump directly to relevant documents.
Implementation Method 1: Inline Citation in Prompt
String prompt = """
根据以下资料回答问题。
资料:
[1] %s
[2] %s
[3] %s
问题:%s
要求:
1. 回答时标注引用的资料编号,格式如 [1][2]
2. 如果多个资料支持同一观点,全部标注
3. 如果资料中没有答案,回答"根据现有资料无法回答"
回答:
""".formatted(doc1, doc2, doc3, question);Output example:
产品A享有2年免费保修服务[1]。保修期内,如出现非人为损坏,
可免费维修或更换[2]。保修期外,维修费用由用户承担[3]。
来源:
[1] 产品手册.pdf - 第3页
[2] 售后政策.pdf - 第1页
[3] 产品手册.pdf - 第4页Implementation Method 2: Structured JSON Output
@Data
public class RagAnswer {
private String answer;
private List<Source> sources;
}
@Data
public class Source {
private String id;
private String title;
private String excerpt;
} String prompt = """
根据资料回答问题,输出 JSON 格式:
{
"answer": "回答内容",
"sources": [
{"id": "1", "title": "来源标题", "excerpt": "引用片段"}
]
}
资料:
%s
问题:%s
""".formatted(context, question);
String jsonResult = chatClient.prompt(prompt).call().content();
RagAnswer answer = objectMapper.readValue(jsonResult, RagAnswer.class);Multi-Document Concatenation Strategies
Simple Concatenation
String context = documents.stream()
.map(Document::getText)
.collect(Collectors.joining("
---
"));Problem: model cannot distinguish which content comes from which document.
Numbered Concatenation
StringBuilder context = new StringBuilder();
for (int i = 0; i < documents.size(); i++) {
Document doc = documents.get(i);
context.append("【资料").append(i + 1).append("】
")
.append(doc.getText())
.append("
");
}Output:
【资料1】
产品A享有2年免费保修服务...
【资料2】
保修期内,如出现非人为损坏...
【资料3】
保修期外,维修费用由用户承担...Metadata-Enriched Concatenation
StringBuilder context = new StringBuilder();
for (int i = 0; i < documents.size(); i++) {
Document doc = documents.get(i);
Map<String, Object> metadata = doc.getMetadata();
context.append("【资料").append(i + 1).append("】
")
.append("来源:").append(metadata.get("source")).append("
")
.append("标题:").append(metadata.get("title")).append("
")
.append("内容:").append(doc.getText()).append("
");
}RAG in Multi-Turn Conversations
Context Merging with Coreference Resolution
public String chatWithRag(String sessionId, String userInput) {
// 1. Get history
List<Message> history = sessionService.getHistory(sessionId);
// 2. Query resolution
String resolvedQuery = coreferenceResolver.resolve(userInput, history);
// 3. Retrieve
List<Document> docs = retriever.retrieve(resolvedQuery);
// 4. Build Prompt
StringBuilder prompt = new StringBuilder();
prompt.append("## 历史对话
");
for (Message msg : history) {
prompt.append(msg.getMessageType()).append(": ")
.append(msg.getText()).append("
");
}
prompt.append("
## 参考资料
");
for (int i = 0; i < docs.size(); i++) {
prompt.append("【资料").append(i + 1).append("】
")
.append(docs.get(i).getText()).append("
");
}
prompt.append("
## 当前问题
").append(userInput);
prompt.append("
## 回答要求
");
prompt.append("基于参考资料回答,如果资料中没有答案,明确告知。");
// 5. Call model
String response = chatClient.prompt(prompt.toString()).call().content();
// 6. Save history
sessionService.addMessage(sessionId, "user", userInput);
sessionService.addMessage(sessionId, "assistant", response);
return response;
}History Compression to Save Tokens
public List<Message> compressHistory(List<Message> history) {
if (history.size() <= 6) return history;
// Keep last 3 rounds (6 messages)
List<Message> recent = history.subList(history.size() - 6, history.size());
// Summarize earlier conversation
String oldText = history.subList(0, history.size() - 6).stream()
.map(m -> m.getMessageType() + ": " + m.getText())
.collect(Collectors.joining("
"));
String summary = chatClient.prompt(
"将以下对话压缩为 100 字以内的摘要:
" + oldText
).call().content();
List<Message> compressed = new ArrayList<>();
compressed.add(new SystemMessage("历史对话摘要:" + summary));
compressed.addAll(recent);
return compressed;
}Complete RAG Prompt Template
@Service
public class RagPromptBuilder {
private static final String SYSTEM_PROMPT = """
你是一个基于企业知识库的问答助手。
## 核心原则
1. 只使用提供的参考资料回答问题
2. 如果资料中没有答案,明确告知用户
3. 回答要准确、简洁、有条理
4. 引用资料时标注编号
""";
private static final String USER_PROMPT = """
## 参考资料
{context}
## 用户问题
{question}
## 回答要求
1. 基于参考资料回答,不要编造
2. 如果资料不足,回答"根据现有资料无法回答该问题"
3. 引用时标注资料编号,如 [1][2]
4. 回答控制在 300 字以内
## 回答
""";
public String build(String question, List<Document> documents) {
// Build numbered context
StringBuilder context = new StringBuilder();
for (int i = 0; i < documents.size(); i++) {
Document doc = documents.get(i);
context.append("【资料").append(i + 1).append("】
")
.append(doc.getText()).append("
");
}
return USER_PROMPT
.replace("{context}", context.toString())
.replace("{question}", question);
}
}Prompt Tuning Checklist
RAG Prompt 调优清单:
约束:
- [ ] 是否明确要求"只基于资料"?
- [ ] 是否处理了"找不到"的情况?
- [ ] 是否禁止模型编造?
格式:
- [ ] 是否要求结构化输出?
- [ ] 是否要求引用来源?
- [ ] 是否限制了回答长度?
上下文:
- [ ] 是否标注了资料编号?
- [ ] 是否包含了来源信息?
- [ ] 是否处理了多文档拼接?
多轮对话:
- [ ] 是否包含了历史对话?
- [ ] 是否做了历史压缩?
- [ ] 是否处理了指代消解?Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
Coder Trainee
Experienced in Java and Python, we share and learn together. For submissions or collaborations, DM us.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
