Agentic RAG: Turning Passive Retrieval into Active Knowledge Agents

This article explains how to combine RAG with AI agents to create Agentic RAG, covering four collaboration patterns (RAG as Tool, Agent as Router, Multi-step RAG, Self-RAG), practical Java Spring Boot implementation with retrieval tools, self-evaluation loops, multi-step reasoning, architecture diagrams, and practical tips on iteration limits and cost monitoring.

Coder Trainee
Coder Trainee
Coder Trainee
Agentic RAG: Turning Passive Retrieval into Active Knowledge Agents

1. The Passive Nature of Traditional RAG

Traditional RAG follows a fixed pipeline:

User Question → Retrieval → Prompt Assembly → Generation → Response

. This flow is hardcoded; regardless of question complexity, the same single retrieval step is executed.

1.1 Three Concrete Limitations

No follow-up questions – Example: "How long is the warranty?" The system does not know which product, yet it never asks for clarification.

No multi-step retrieval – Example: "What is the difference between A and B?" A single retrieval cannot fetch complete information for both entities.

No judgment on retrieval necessity – Example: "How is the weather today?" The knowledge base lacks this data, but the system still performs a retrieval.

Core issue: Traditional RAG lacks decision-making capability.

2. RAG + Agent Collaboration Patterns

2.1 From Passive Retrieval to Active Retrieval

┌─────────────────────────────────────────────────────────────────┐
│         Traditional RAG vs Agentic RAG                          │
├─────────────────────────────────────────────────────────────────┤
│                                                                 │
│ Traditional RAG:                                                │
│ Question → Retrieval → Generation → Answer                      │
│ (Single retrieval, fixed flow)                                  │
│                                                                 │
│ Agentic RAG:                                                    │
│ Question → Agent Reasoning → Decide Whether to Retrieve → Retrieve → Evaluate Results │
│                              │                    │
│                              │                    ├── Not enough? → Rewrite query and retrieve again │
│                              │                    │
│                              │                    └── Enough → Generate answer │
│                              │
│                              └── No retrieval needed? → Answer directly │
│                                                                 │
└─────────────────────────────────────────────────────────────────┘

2.2 Four Collaboration Modes

RAG as Tool – Wrap retrieval as a tool; the agent calls it on demand. Suitable for general scenarios.

Agent as Router – Agent decides which knowledge base to query. Suitable for multiple knowledge bases.

Multi-step RAG – Agent performs multiple retrievals, iteratively approaching the answer. Suitable for complex questions.

Self-RAG – Agent evaluates its own generated answer. Suitable for high-quality requirements.

3. Practical Implementation: Wrapping Retrieval as an Agent Tool

3.1 Defining Retrieval Tools (Java Spring)

@Component
public class RetrievalTools {

    @Autowired
    private HybridRetriever hybridRetriever;

    @Autowired
    private Reranker reranker;

    @Tool(description = "Retrieve information from the knowledge base. Use when looking up enterprise documents, product manuals, policy regulations.")
    public String searchKnowledge(
        @ToolParam(description = "Search query, should be a clear, specific question") String query) {

        List<Document> candidates = hybridRetriever.retrieve(query, 20);
        List<Document> reranked = reranker.rerank(query, candidates, 5);

        if (reranked.isEmpty()) {
            return "No relevant information found";
        }

        StringBuilder result = new StringBuilder();
        for (int i = 0; i < reranked.size(); i++) {
            result.append("[Source " + (i + 1) + "]
")
                  .append(reranked.get(i).getText()).append("

");
        }
        return result.toString();
    }

    @Tool(description = "Get full content of a specified document. Use when retrieved snippets are incomplete.")
    public String getDocumentDetail(
        @ToolParam(description = "Document ID") String documentId) {
        // Fetch full document from database
        return documentRepository.findById(documentId)
                .map(DocumentEntity::getContent)
                .orElse("Document not found");
    }
}

3.2 Agent Configuration

@Configuration
public class RagAgentConfig {

    @Bean
    public ChatClient ragAgent(ChatClient.Builder builder, RetrievalTools tools) {
        return builder
            .defaultSystem("""
                You are an enterprise knowledge base assistant.

                ## Core Principles
                1. Prefer using searchKnowledge tool to find information
                2. If retrieved information is incomplete, retrieve again or get document details
                3. If knowledge base truly lacks relevant information, explicitly tell the user
                4. Cite sources in answers

                ## Workflow
                1. Understand user question
                2. Decide whether retrieval is needed
                3. If needed, call searchKnowledge
                4. Evaluate whether retrieval results are sufficient
                5. If not, adjust query and retrieve again
                6. Generate final answer
                """)
            .defaultTools(tools)
            .build();
    }
}

3.3 Usage

@RestController
public class RagAgentController {

    @Autowired
    private ChatClient ragAgent;

    @PostMapping("/chat")
    public String chat(@RequestBody String question) {
        return ragAgent.prompt(question).call().content();
    }
}

The agent automatically decides:

Does this question require retrieval?

Are the retrieval results sufficient?

Is another retrieval needed?

4. Self-RAG: Agent Evaluates Its Own Answers

4.1 Core Idea

Self-RAG's core: after generating an answer, let the agent evaluate whether the answer is grounded in the retrieved content; if not good enough, re-retrieve.

@Service
public class SelfRagService {

    private final ChatClient agent;

    public String query(String question) {
        int maxAttempts = 3;

        for (int attempt = 0; attempt < maxAttempts; attempt++) {
            // 1. Retrieve
            List<Document> docs = retrieve(question);

            // 2. Generate
            String answer = generate(question, docs);

            // 3. Self-evaluation
            Evaluation eval = selfEvaluate(question, answer, docs);

            if (eval.isGood()) {
                return answer;
            }

            // 4. If not good enough, rewrite query and retry
            question = rewriteQuery(question, eval.getFeedback());
            log.info("Attempt {}, evaluation result: {}", attempt + 1, eval.getFeedback());
        }

        return "Sorry, I cannot accurately answer this question.";
    }

    private Evaluation selfEvaluate(String question, String answer, List<Document> docs) {
        String context = docs.stream()
                .map(Document::getText)
                .collect(Collectors.joining("
"));

        String prompt = """
            Evaluate whether the following answer is based on the provided sources.

            Sources: %s

            Question: %s

            Answer: %s

            Evaluation:
            1. Is the answer fully based on sources? (yes/no)
            2. Does the answer sufficiently address the question? (yes/no)
            3. If incomplete, how should it be improved?

            Output JSON: {"good": true/false, "feedback": "..."}
            """.formatted(context, question, answer);

        String result = agent.prompt(prompt).call().content();
        return parseEvaluation(result);
    }
}

5. Multi-step Reasoning RAG

5.1 Decomposing Complex Questions

User asks: "Compared to product A, which is more suitable for enterprise users, product A or product B?"

Answering this requires:

Product A's positioning

Product B's positioning

Enterprise user requirements

Comparative analysis

@Service
public class MultiStepRagService {

    public String query(String question) {
        // 1. Agent decomposes the question
        List<String> subQuestions = decompose(question);
        // ["Product A positioning and features", "Product B positioning and features", "Enterprise user requirements"]

        // 2. Retrieve each sub-question
        Map<String, String> results = new HashMap<>();
        for (String subQ : subQuestions) {
            List<Document> docs = retrieve(subQ);
            results.put(subQ, formatDocs(docs));
        }

        // 3. Synthesize final answer
        String prompt = buildSynthesisPrompt(question, results);
        return generate(prompt);
    }
}

6. Complete Agentic RAG Architecture

┌─────────────────────────────────────────────────────────────────┐
│                     Agentic RAG Architecture                    │
├─────────────────────────────────────────────────────────────────┤
│                                                                 │
│ User Question                                                   │
│      │                                                          │
│      ▼                                                          │
│ ┌─────────────────────────────────────────────────────────┐   │
│ │                      Agent                                │   │
│ │                                                         │   │
│ │ 1. Analyze Question                                     │   │
│ │ 2. Decide Whether Retrieval Needed                      │   │
│ │ 3. Choose Retrieval Strategy                            │   │
│ │ 4. Call Retrieval Tools                                 │   │
│ │ 5. Evaluate Results                                     │   │
│ │ 6. Decide Whether to Continue                           │   │
│ │ 7. Generate Answer                                      │   │
│ └─────────────────────────────────────────────────────────┘   │
│      │                                                          │
│      ├── Retrieval Tools → Vector Store / BM25 / Graph        │
│      ├── Document Tools → Full Document Fetch                 │
│      ├── Compute Tools → Mathematical Calculations            │
│      └── Other Tools → Business Systems                       │
│                                                                 │
└─────────────────────────────────────────────────────────────────┘

7. Practical Recommendations

Start Simple – Begin with RAG as Tool; only add Self-RAG if needed.

Control Iteration Count – Set a maximum number of iterations (e.g., 3).

Monitor Costs – Agentic RAG token consumption is 3-5 times that of traditional RAG.

Evaluate Effectiveness – Compare traditional RAG and Agentic RAG performance differences.

8. Next Episode Preview

Building AI Knowledge Base from Scratch (16): Series Finale – The Future Evolution of RAG

Technology trends

From RAG to Agentic RAG

Learning roadmap

I am Lao J, see you next time.

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

AI AgentsRAGSpring BootKnowledge BaseRetrieval-Augmented GenerationAgentic RAGSelf-RAGMulti-step RAG
Coder Trainee
Written by

Coder Trainee

Experienced in Java and Python, we share and learn together. For submissions or collaborations, DM us.

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.