Build a Production-Ready Enterprise RAG Knowledge Base: Full Project Walkthrough

This final installment of a 12-part RAG series demonstrates a complete, deployable enterprise knowledge base system using Spring Boot 3.2, Spring AI 2.0, PgVector, BGE embeddings, hybrid retrieval with RRF fusion, reranking, multi-turn chat with memory, and a Vue 3 frontend, including Docker deployment and production checklist.

Coder Trainee
Coder Trainee
Coder Trainee
Build a Production-Ready Enterprise RAG Knowledge Base: Full Project Walkthrough

Project Overview

Goals

Build an enterprise knowledge base QA system supporting:

Multi-format document import (PDF/Word/HTML/Markdown)

Intelligent retrieval (hybrid search + reranking)

Multi-turn conversation (memory + query optimization)

Source citations

Admin dashboard (document management, evaluation monitoring)

Tech Stack

Framework: Spring Boot 3.2 + Spring AI 2.0

Vector Store: PgVector

Embedding: BGE-small-zh-v1.5 (local)

Reranker: BGE-Reranker-base (local)

Cache: Redis + Caffeine

Database: PostgreSQL

Frontend: Vue 3 + Element Plus

Project Structure

enterprise-knowledge-base/
├── pom.xml
├── src/main/java/com/example/kb/
│   ├── KnowledgeBaseApplication.java
│   ├── config/
│   │   ├── AiConfig.java
│   │   ├── VectorStoreConfig.java
│   │   └── RedisConfig.java
│   ├── controller/
│   │   ├── ChatController.java
│   │   ├── DocumentController.java
│   │   └── AdminController.java
│   ├── service/
│   │   ├── ingestion/
│   │   │   ├── DocumentIngestionService.java
│   │   │   ├── DocumentParser.java
│   │   │   └── ChunkingService.java
│   │   ├── retrieval/
│   │   │   ├── HybridRetriever.java
│   │   │   ├── Reranker.java
│   │   │   └── QueryOptimizer.java
│   │   ├── generation/
│   │   │   ├── RagPromptBuilder.java
│   │   │   └── AnswerGenerator.java
│   │   └── evaluation/
│   │       └── EvaluationService.java
│   ├── model/
│   │   ├── ChatRequest.java
│   │   ├── ChatResponse.java
│   │   └── Document.java
│   └── repository/
│       └── DocumentRepository.java
└── src/main/resources/
    ├── application.yml
    └── prompts/
        └── rag-system.md

Core Implementation

Maven Dependencies

<dependencies>
  <!-- Spring AI -->
  <dependency>
    <groupId>org.springframework.ai</groupId>
    <artifactId>spring-ai-openai-spring-boot-starter</artifactId>
    <version>2.0.0</version>
  </dependency>
  <dependency>
    <groupId>org.springframework.ai</groupId>
    <artifactId>spring-ai-pgvector-store-spring-boot-starter</artifactId>
    <version>2.0.0</version>
  </dependency>
  <!-- Document Parsing -->
  <dependency>
    <groupId>org.apache.pdfbox</groupId>
    <artifactId>pdfbox</artifactId>
    <version>3.0.0</version>
  </dependency>
  <dependency>
    <groupId>org.apache.poi</groupId>
    <artifactId>poi-ooxml</artifactId>
    <version>5.2.5</version>
  </dependency>
  <dependency>
    <groupId>org.jsoup</groupId>
    <artifactId>jsoup</artifactId>
    <version>1.17.2</version>
  </dependency>
  <!-- Web -->
  <dependency>
    <groupId>org.springframework.boot</groupId>
    <artifactId>spring-boot-starter-web</artifactId>
  </dependency>
  <dependency>
    <groupId>org.springframework.boot</groupId>
    <artifactId>spring-boot-starter-data-redis</artifactId>
  </dependency>
</dependencies>

Configuration (application.yml)

spring:
  application:
    name: enterprise-knowledge-base
  datasource:
    url: jdbc:postgresql://localhost:5432/knowledge
    username: postgres
    password: ${DB_PASSWORD}
  ai:
    openai:
      api-key: ${OPENAI_API_KEY}
      chat:
        options:
          model: gpt-4
          temperature: 0.3
      vectorstore:
        pgvector:
          initialize-schema: true
          dimensions: 512
          distance-type: COSINE
  data:
    redis:
      host: localhost
      port: 6379
server:
  port: 8080
rag:
  chunk-size: 500
  chunk-overlap: 100
  retrieval-top-k: 20
  rerank-top-k: 5
  similarity-threshold: 0.7

Document Ingestion Service

@Service
@Slf4j
public class DocumentIngestionService {
    @Autowired private DocumentParser parser;
    @Autowired private ChunkingService chunkingService;
    @Autowired private VectorStore vectorStore;
    @Autowired private DocumentRepository documentRepository;

    @Async("documentExecutor")
    public CompletableFuture<IngestionResult> ingest(MultipartFile file) {
        try {
            // 1. Save raw file info
            DocumentEntity entity = saveDocumentEntity(file);
            // 2. Parse
            String content = parser.parse(file);
            log.info("Document parsed: {} chars", content.length());
            // 3. Chunk
            List<Chunk> chunks = chunkingService.chunk(content, entity);
            // 4. Embedding + store
            List<org.springframework.ai.document.Document> aiDocs = chunks.stream()
                .map(this::toAiDocument)
                .collect(Collectors.toList());
            vectorStore.add(aiDocs);
            log.info("Vector storage completed: {} chunks", aiDocs.size());
            // 5. Update status
            entity.setStatus("COMPLETED");
            entity.setChunkCount(chunks.size());
            documentRepository.save(entity);
            return CompletableFuture.completedFuture(
                IngestionResult.success(entity.getId(), chunks.size()));
        } catch (Exception e) {
            log.error("Document ingestion failed", e);
            return CompletableFuture.failedFuture(e);
        }
    }

    private org.springframework.ai.document.Document toAiDocument(Chunk chunk) {
        return new org.springframework.ai.document.Document(
            chunk.getContent(),
            Map.of(
                "document_id", chunk.getDocumentId(),
                "title", chunk.getTitle(),
                "section", chunk.getSection(),
                "chunk_index", chunk.getIndex()
            )
        );
    }
}

Hybrid Retrieval Service

@Service
public class HybridRetriever {
    @Autowired private VectorStore vectorStore;
    @Autowired private BM25Retriever bm25Retriever;
    private static final int RRF_K = 60;

    public List<Document> retrieve(String query, int topK) {
        // 1. Vector search
        List<Document> vectorResults = vectorStore.similaritySearch(
            SearchRequest.builder()
                .query(query)
                .topK(topK * 2)
                .similarityThreshold(0.6)
                .build()
        );
        // 2. BM25 search
        List<Document> keywordResults = bm25Retriever.search(query, topK * 2);
        // 3. RRF fusion
        return rrfFusion(vectorResults, keywordResults, topK);
    }

    private List<Document> rrfFusion(List<Document> list1, List<Document> list2, int topK) {
        Map<String, Double> scores = new HashMap<>();
        Map<String, Document> docMap = new HashMap<>();
        for (int i = 0; i < list1.size(); i++) {
            String id = list1.get(i).getId();
            scores.merge(id, 1.0 / (RRF_K + i + 1), Double::sum);
            docMap.put(id, list1.get(i));
        }
        for (int i = 0; i < list2.size(); i++) {
            String id = list2.get(i).getId();
            scores.merge(id, 1.0 / (RRF_K + i + 1), Double::sum);
            docMap.putIfAbsent(id, list2.get(i));
        }
        return scores.entrySet().stream()
            .sorted(Map.Entry.<String, Double>comparingByValue().reversed())
            .limit(topK)
            .map(e -> docMap.get(e.getKey()))
            .collect(Collectors.toList());
    }
}

RAG Chat Service

@Service
@Slf4j
public class RagChatService {
    @Autowired private QueryOptimizer queryOptimizer;
    @Autowired private HybridRetriever retriever;
    @Autowired private Reranker reranker;
    @Autowired private RagPromptBuilder promptBuilder;
    @Autowired private ChatClient chatClient;
    @Autowired private SessionService sessionService;

    public ChatResponse chat(String sessionId, String userInput) {
        long start = System.currentTimeMillis();
        // 1. Get history + optimize query
        List<Message> history = sessionService.getHistory(sessionId);
        String optimizedQuery = queryOptimizer.optimize(userInput, history);
        // 2. Hybrid retrieval
        List<Document> candidates = retriever.retrieve(optimizedQuery, 20);
        // 3. Rerank
        List<Document> reranked = reranker.rerank(optimizedQuery, candidates, 5);
        // 4. Build prompt
        String prompt = promptBuilder.build(userInput, reranked);
        // 5. Generate answer
        String answer = chatClient.prompt(prompt).call().content();
        // 6. Save history
        sessionService.addMessage(sessionId, "user", userInput);
        sessionService.addMessage(sessionId, "assistant", answer);
        long duration = System.currentTimeMillis() - start;
        return ChatResponse.builder()
            .answer(answer)
            .sources(extractSources(reranked))
            .durationMs(duration)
            .build();
    }

    private List<Source> extractSources(List<Document> documents) {
        return documents.stream()
            .map(doc -> Source.builder()
                .documentId(doc.getMetadata().get("document_id").toString())
                .title(doc.getMetadata().get("title").toString())
                .section(doc.getMetadata().get("section").toString())
                .excerpt(doc.getText().substring(0, Math.min(100, doc.getText().length())))
                .build())
            .collect(Collectors.toList());
    }
}

Controllers

@RestController
@RequestMapping("/api")
public class ChatController {
    @Autowired private RagChatService chatService;

    @PostMapping("/chat")
    public ChatResponse chat(@RequestBody ChatRequest request) {
        String sessionId = request.getSessionId() != null
            ? request.getSessionId()
            : UUID.randomUUID().toString();
        return chatService.chat(sessionId, request.getMessage());
    }

    @PostMapping("/chat/stream")
    public Flux<String> chatStream(@RequestBody ChatRequest request) {
        return chatService.chatStream(request.getSessionId(), request.getMessage());
    }
}

@RestController
@RequestMapping("/api/documents")
public class DocumentController {
    @Autowired private DocumentIngestionService ingestionService;

    @PostMapping("/upload")
    public Map<String, Object> upload(@RequestParam("file") MultipartFile file) {
        CompletableFuture<IngestionResult> future = ingestionService.ingest(file);
        return Map.of(
            "status", "PROCESSING",
            "message", "Document submitted, processing in background"
        );
    }

    @GetMapping("/list")
    public List<DocumentEntity> list() {
        return documentRepository.findAll();
    }

    @DeleteMapping("/{id}")
    public Map<String, Object> delete(@PathVariable Long id) {
        ingestionService.deleteDocument(id);
        return Map.of("status", "DELETED");
    }
}

Prompt Template (prompts/rag-system.md)

You are a QA assistant based on an enterprise knowledge base.

## Core Principles
1. Only use provided reference materials to answer
2. If materials lack the answer, explicitly tell the user
3. Cite sources with numbers
4. Be accurate, concise, organized

## Reference Materials
{context}

## User Question
{question}

## Answer Requirements
1. Answer based on references, do not fabricate
2. If insufficient, reply "Unable to answer based on available materials"
3. Cite with numbers like [1][2]
4. Keep answer under 300 words

## Answer

Frontend Interface

Core Features

┌─────────────────────────────────────────────────────────────────┐
│              Enterprise Knowledge Base UI                       │
├─────────────────────────────────────────────────────────────────┤
│                                                                 │
│  Sidebar                          │  Main Area                  │
│  ├── Chat                         │  ├── Chat Window            │
│  ├── Document Management          │  ├── Message List           │
│  ├── Knowledge Base Settings      │  ├── Input Box              │
│  └── Evaluation Reports           │  └── Source Display         │
│                                                                 │
└─────────────────────────────────────────────────────────────────┘

Chat Component (Vue 3)

<template>
  <div class="chat-container">
    <div class="messages">
      <div v-for="msg in messages" :key="msg.id" :class="['message', msg.role]">
        <div class="content">{{ msg.content }}</div>
        <div v-if="msg.sources" class="sources">
          <div class="source-label">Sources:</div>
          <div v-for="src in msg.sources" :key="src.id" class="source-item">
            <span class="source-title">{{ src.title }}</span>
            <span class="source-section">{{ src.section }}</span>
          </div>
        </div>
      </div>
    </div>
    <div class="input-area">
      <el-input v-model="input" @keyup.enter="send" placeholder="Enter question..." />
      <el-button @click="send">Send</el-button>
    </div>
  </div>
</template>

Deployment

Docker Compose

version: '3.8'
services:
  postgres:
    image: pgvector/pgvector:pg16
    environment:
      POSTGRES_DB: knowledge
      POSTGRES_PASSWORD: ${DB_PASSWORD}
    volumes:
      - pgdata:/var/lib/postgresql/data
    ports:
      - "5432:5432"

  redis:
    image: redis:7-alpine
    ports:
      - "6379:6379"

  kb-app:
    build: .
    environment:
      - DB_PASSWORD=${DB_PASSWORD}
      - OPENAI_API_KEY=${OPENAI_API_KEY}
    ports:
      - "8080:8080"
    depends_on:
      - postgres
      - redis

volumes:
  pgdata:

Production Checklist

Infrastructure:

PostgreSQL (primary-replica)

Redis (Sentinel/Cluster)

Vector Database (PgVector/Milvus)

Application:

Multi-replica deployment

Health checks

Graceful shutdown

Resource limits

Security:

HTTPS

API Key management

Rate limiting

Access control

Monitoring:

Application metrics

Business metrics

Alerting

Log aggregation

Series Summary

Twelve episodes covering the full RAG lifecycle from document parsing to production deployment:

Episode 1: Why RAG Performs Poorly

Episode 2: Document Parsing & Cleaning

Episode 3: Chunking Strategies

Episode 4: Embedding Model Selection

Episode 5: Vector Database Selection

Episode 6: Hybrid Retrieval

Episode 7: Reranking

Episode 8: Query Optimization

Episode 9: Prompt Engineering

Episode 10: RAG Evaluation Framework

Episode 11: Production-Grade Architecture

Episode 12: Hands-on Project

RAG is not "a feature" but a system. Every component deserves careful attention.

This concludes the RAG series. See you in the next series.

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

DockerRAGSpring BootSpring AIVue 3Hybrid RetrievalPgVectorBGERerankingEnterprise Knowledge Base
Coder Trainee
Written by

Coder Trainee

Experienced in Java and Python, we share and learn together. For submissions or collaborations, DM us.

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.