RAG Deep Dive: Principles, Code & Engineering Practices for Retrieval-Augmented Generation

This article provides a comprehensive technical analysis of Retrieval-Augmented Generation (RAG), covering embedding principles, chunking strategies, vector databases, dense/sparse/hybrid retrieval, reranking, HyDE, GraphRAG, and engineering evaluation frameworks with code examples.

Tech Bean
Tech Bean
Tech Bean
RAG Deep Dive: Principles, Code & Engineering Practices for Retrieval-Augmented Generation

Retrieval-Augmented Generation (RAG) addresses three structural defects of large language models: knowledge cutoff, hallucination, and missing private knowledge. RAG formalizes the problem as: given query q, retrieve Top-K relevant chunks R=TopK(Retrieve(q,D)) from knowledge base D, then generate answer conditioned on (q, R). Unlike fine-tuning, RAG does not modify model parameters, offering instant updates, low cost, and traceability.

Embedding Principles and Vectorization

Embeddings map unstructured text to dense vectors (e.g., 768/1024/1536/3072 dimensions) so semantically similar texts are closer in vector space. Models are trained via contrastive learning (InfoNCE, in-batch negative sampling) to pull positive pairs together and push negatives apart. The article provides a minimal sentence-transformers example using BAAI/bge-large-zh-v1.5 to produce a 1024-dim vector.

RAG Pipeline

Two pipelines: Indexing (offline) – documents → cleaning/chunking → embedding → vector store with metadata; Query (online) → same embedding → similarity search → Top-K chunks → assemble prompt → LLM generation.

Document Chunking Strategies

Fixed-length chunking : hard cut by token count (300–800 tokens, 50–150 overlap), simple but breaks semantic boundaries.

Recursive chunking : hierarchical separators (paragraph → sentence → phrase → word), engineering default.

Structural chunking : uses Markdown headings/sections, best traceability.

Semantic chunking : computes embedding similarity between adjacent sentence groups; sharp drops indicate topic shifts.

Parent-child chunking : small child chunks for high-precision recall, then return parent chunk for full context.

Code example shows RecursiveCharacterTextSplitter with chunk_size=512, chunk_overlap=80, and Chinese punctuation separators.

PDF Extraction Pitfalls

Headers/footers/page numbers and line-break hyphenation → filter by layout coordinates, reassemble by paragraph.

Scanned PDFs (image-only) → OCR with tools like Marker to output Markdown.

Hundreds of pages → streaming page-by-page reading to avoid memory overload.

Table-heavy PDFs → plain text extraction destroys row/column relations; prefer LlamaParse to Markdown tables as independent chunks.

Entity inconsistency across chunks (e.g., “LLM” vs “大语言模型”) → entity resolution to unify nodes.

Embedding Model Selection and Evaluation

Seven dimensions: vector dimension (storage/compute cost), recall quality (MTEB leaderboard), inference speed, VRAM usage, language support (Chinese), max input length (must match chunk_size), domain adaptation. Evaluation benchmark: MTEB (HuggingFace) covering multiple tasks; Chinese scenarios use C-MTEB with 6 task types, 35 datasets across e-commerce, medical, policy, news, search. Leaderboard: huggingface.co/spaces/mteb/leaderboard. Minimal evaluation code uses mteb library with MTEB-Chinese task set. Common choices: OpenAI text-embedding-3-small/large (hosted, small cheaper, large better); open-source BGE/GTE/Qwen-Embedding (private deployment, C-MTEB validated, dynamic quantization + ONNX acceleration). Critical note: a knowledge base must fix a single embedding model; any change requires full index rebuild.

Vector Databases and ANN Indexes

Each record: vector (embedding), text (original chunk), metadata (source, page, time, permissions). Production uses Approximate Nearest Neighbor (ANN) indexes, not brute-force O(N×d):

HNSW : multi-layer graph, coarse-to-fine search, ~O(log N) query, balances recall and latency.

IVF : clusters vector space (e.g., 4096 cells), scans only hit cells, suited for massive libraries, recall controlled by nprobe.

Comparison table of seven databases: Qdrant (open-source + cloud, Docker/cluster, HNSW, sharding+replication, Rust low-latency, RAG/production), Milvus (open-source + Zilliz, storage-compute separation, HNSW/IVF/DiskANN, trillion-scale, unstructured data platform), Chroma (open-source, embedded/local file, HNSW via hnswlib, single-node lightweight, prototyping/teaching/small RAG), Weaviate (open-source, Docker/K8s, HNSW, native hybrid BM25+vector, semantic search/multimodal), Pinecone (commercial SaaS, fully managed, hierarchical index, zero-ops, pay-as-you-go, rapid production), pgvector (PostgreSQL extension, IVFFlat/HNSW, scales with PG, existing PG + moderate vector scale), Faiss (Meta library, in-process, IVF/HNSW/PQ, algorithmic building block, research/custom index/offline retrieval).

Qdrant Deployment and Usage

Development: Docker single-node mapping ports 6333 (REST) and 6334 (cluster), volume mount for persistence. Health checks: GET /readyz, GET /livez, dashboard at /dashboard. Production: K8s StatefulSet multi-replica, snapshot backups, shard by collection for horizontal scaling. Code snippets show collection creation (1024-dim, cosine), upsert with payload (text, source, page), basic Top-K query, and metadata-filtered query using Filter and FieldCondition.

Retrieval: Dense, Sparse, Hybrid

Dense Retrieval

Embed query, compute cosine similarity with document vectors, return Top-K. Strong semantic generalization, weak on exact identifiers (model numbers, codes, proper nouns). In-memory simulation code demonstrates filtering, batch embedding, cosine scoring, sorting, and Top-K cutoff.

Sparse Retrieval

Converts text to high-dimensional sparse vectors (vocabulary size ~hundreds of thousands). Only non-zero dimensions for terms present in document. Relies on keywords, term frequency, inverted index. Key concepts: TF (term frequency in document), DF (document frequency), IDF (inverse document frequency, rare terms get higher weight), Posting List mapping term → [(doc_id, tf), ...], score aggregation by summing term scores across query terms. BM25 scoring formula shown (image). L2 normalization explained with example (image). Dot product computation shown (image). Engineering implementation: sparse index must be built at indexing time. Two backend paths: (1) stores with native BM25 (Elasticsearch) – TF/IDF maintained automatically; (2) vector databases like Qdrant – client must tokenize, compute IDF, generate sparse vectors (e.g., SPLADE) at write time, store as sparse vector field alongside dense vector, retrieve via sparse dot product, then fuse. Full Haystack-style code shows fit() building vocabulary and smoothed IDF idf(t) = log((1+N)/(1+df(t))) + 1, _to_sparse() producing TF-IDF sparse embedding with normalized TF, sorted indices for Qdrant.

Hybrid Search

Combines keyword and vector retrieval, fuses results for final ranking. Fusion methods:

Weighted reranking : normalize BM25 and vector scores, apply weights (e.g., 0.4 * S_bm25_norm + 0.6 * S_vec_norm), sort by combined score.

RRF (Reciprocal Rank Fusion) : discards raw scores, uses rank positions only: RRF(d) = Σ 1/(k + rank(d)) with k=60. Example shows sparse ranks A=1,B=2,C=3,D=4 and dense ranks B=1,E=2,A=3,F=4 → RRF(A)=1/61+1/63.

Reranker (Cross-Encoder)

Takes [query, candidate_doc] pairs, outputs 0–1 relevance score, reorders top candidates. Needed because hybrid fusion is still coarse: independent encoders cannot see each other, leading to false positives (topically similar but not answering). Reranker jointly encodes query+doc for precise judgment. Trade-off: slower, higher compute; only applied to small candidate set (Top-30~50 from coarse retrieval). Code uses BAAI/bge-reranker-base via CrossEncoder, predicts on pairs, sorts descending, takes Top-5.

HyDE: Hypothetical Document Embeddings

LLM generates a hypothetical answer document from query, then embeds that document for retrieval. Short queries have sparse semantics; hypothetical document better matches corpus distribution, boosting recall. Cost: extra LLM call and potential bias. Example:

hyde_doc = llm.invoke(f"请撰写一段可能回答该问题的文档:{query}")

then query_vec = embed_model.encode(hyde_doc).

Graph RAG: Graph-Structured Retrieval

Traditional RAG chunks knowledge into independent vectors, breaking on cross-document, multi-hop reasoning, and global summarization. Graph RAG introduces knowledge graph: LLM extracts (head, relation, tail) triplets (S-P-O) from unstructured text. Index construction in five steps:

TextUnit chunking : chunk parameters tuned for stable triplet extraction.

Triplet extraction (S-P-O) : LLM per-chunk extracts entities and relations, e.g., (刘强东, 创立, 京东). Code uses LLMGraphTransformer with gpt-4o-mini to produce nodes and relationships.

Leiden community clustering : clusters triplet edges into thematic communities, forming hierarchical knowledge system.

Hierarchical community summaries : LLM generates structured summaries per community (topic, key entities, core relations, key facts, limitations) to reduce hallucination.

Vector storage : embeddings for raw TextUnits, entity descriptions, low-level community summaries, high-level community summaries stored in three separate vector spaces: chunk_index, entity_index, community_index.

Query modes: Local Search (detail/entity questions) – identify key entities → BFS graph traversal → associate raw text → rerank → LLM generate. Global Search (summarization/induction) – skip fine-grained entities, vector-match high-level community summaries, feed multiple high-relevance summaries to LLM for global synthesis. Common post-step: retrieved subgraph + text chunks → reranker (BGE-Reranker) → filter low relevance → LLM final answer.

Engineering Practices: Evaluation Metrics and Decision Framework

Pre-production quantitative evaluation system:

Recall@K / Hit Rate : whether ground-truth chunk enters Top-K, measures recall capability.

MRR (Mean Reciprocal Rank) : rank quality of ground-truth in results.

Answer correctness (LLM-as-Judge) : strong model scores “faithfulness (grounded in retrieved content)” and “correctness”.

Four-step engineering decision process:

Scenario determination : introduce RAG only when answers depend on private docs, real-time info, or require citation; pure common-knowledge QA doesn't need retrieval complexity.

Data pipeline first : cleaning (headers/OCR/tables) → chunking strategy (structural/recursive start) → small-sample manual recall verification.

Retrieval pipeline baseline : hybrid search (BM25 + dense + RRF) → coarse Top-50 → reranker precision Top-5.

Evaluation-driven iteration : fixed eval set, single-variable experiments (chunk_size, Top-K, model, rerank on/off), decide by metrics not intuition.

RAG essence: decouples “parametric memory” from “external verifiable knowledge”. Model handles language composition, knowledge base supplies facts, retrieval aligns them. Engineering ceiling set by retrieval quality and data pipeline quality; model choice only sets the floor.

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

RAGBM25Retrieval-Augmented GenerationVector DatabasesEmbeddingsGraphRAGHybrid SearchReranker
Tech Bean
Written by

Tech Bean

Learning, sharing, and news

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.