Beyond RAG: Vector Databases as General-Purpose Similarity Search Engines
This article explains why vector databases are general-purpose similarity search engines beyond RAG, covering embedding model selection (BGE), non-RAG use cases like recommendation and anomaly detection, core operations with ChromaDB and HNSW, selection criteria for ChromaDB, FAISS, Milvus, and Pinecone, and five practical engineering lessons on recall evaluation, API workarounds, distance vs similarity, chunking strategies, and metadata filtering.
Why Vector Databases Are Needed
Traditional keyword search (Elasticsearch, SQL LIKE) matches literal text. A query for "印染废水 COD 限值" only finds documents containing those exact characters, missing synonyms, phrasing variations, or semantic equivalents. Vector search matches semantics: both the query and a document like "纺织染整工业水污染物排放标准规定 COD 排放限值 200mg/L" are embedded into high‑dimensional vectors, and cosine similarity retrieves them even with zero shared terms.
The challenge is scale: exact nearest‑neighbor search is O(Nd) linear scan, feasible for 10k vectors but not 100M. Vector databases solve this with ANN (Approximate Nearest Neighbor) indexes — sacrificing a tiny amount of precision for sub‑linear query time, returning results that are "close enough" rather than mathematically exact.
Embedding: Turning Text into Vectors
Vector databases store float arrays, not text. Embedding models compress discrete text into a continuous high‑dimensional space where semantically similar texts are neighbors. The project uses BGE (BAAI/bge-small-zh-v1.5) for three reasons:
Chinese performance: trained specifically for Chinese, not an English model with added Chinese support.
Small footprint: ~93 MB, runs on CPU without GPU.
Industry adoption: mature community, strong reproducibility.
A critical detail: BGE requires different prefixes for queries and documents. The query side adds an instruction prefix:
emb = self._model.encode(
f"为这个句子生成表示以用于检索相关文章:{text}"
)Documents are encoded without the prefix. Forgetting the prefix drops recall by 5–15 percentage points while the code still runs, making the regression hard to detect. This also illustrates that embedding models differ in how they define similarity — BGE separates query intent from document content, others do not.
Beyond RAG: Other Vector Database Use Cases
Viewing vector databases as general similarity search engines reveals many non‑RAG applications:
Recommendation systems: user history embedding + item embedding → nearest neighbors. Powers "you may like", "similar videos", "related news" without LLMs or complex rule engines.
Document deduplication & anomaly detection: two sides of the same problem. Embed known normal samples; new samples with similarity >0.95 are duplicates (news crawlers skip them), similarity <0.3 are anomalies (industrial IoT flags abnormal logs). Same formula, different thresholds.
Long‑context compression: split a large report or contract into N chunks, embed all, retrieve only the top‑3 relevant chunks for the LLM. The vector database acts as a selector for limited context windows.
Abstracting the vector database as team infrastructure yields more reuse than a one‑line import in a RAG tutorial.
Basic Vector Database Operations
APIs across vector databases share a common skeleton: create index, query, delete. ChromaDB example:
# Create
client = chromadb.PersistentClient(path="./data/chroma_db")
collection = client.get_or_create_collection("env_laws")
# Insert
collection.add(documents=chunks, embeddings=embeddings, metadatas=metadatas, ids=ids)
# Query
results = collection.query(query_embeddings=[query_vec], n_results=5)The core indexing algorithm is HNSW (Hierarchical Navigable Small World) . Analogy: a city transport network — sparse express highways (top layer) quickly bring you to the target district, then dense local streets (bottom layer) find the exact neighbor. Query starts at the top layer, complexity approaches O(log N). Milvus, Qdrant, Weaviate all default to HNSW or variants.
However, selection hinges on engineering traits (deployment model, scalability, ecosystem), not which index algorithm is newest.
Mainstream Vector Database Selection
Practical guidance by scale:
Tens to hundreds of thousands → ChromaDB to get running fast.
Performance‑sensitive, single‑node → FAISS .
Billion‑scale, distributed → Milvus .
Zero‑ops preference → Pinecone (managed service).
The landscape splits into embedded (ChromaDB, FAISS, raw numpy) and service‑oriented (Milvus, Weaviate, Qdrant, Pinecone) architectures.
Five Practical Lessons
1. Evaluation Must Go Beyond "It Runs"
Before launch, measure recall: the fraction of truly relevant documents retrieved. Method: manually annotate 50–100 "question → relevant document" pairs, run retrieval, compute hit rate.
2. Don’t Wrestle with Broken APIs
ChromaDB 0.5.5 clashed with huggingface-hub versions, causing collection.query() errors. Instead of debugging, the team bypassed ChromaDB’s query API: read documents directly from the underlying chroma.sqlite3, then used sentence-transformers + NumPy cosine similarity for retrieval:
# Read directly from ChromaDB SQLite, skip query API
cursor.execute("""
SELECT em.string_value, em2.string_value, em3.string_value
FROM embedding_metadata em
LEFT JOIN embedding_metadata em2 ON em.id = em2.id AND em2.key = 'law_name'
LEFT JOIN embedding_metadata em3 ON em.id = em3.id AND em3.key = 'source'
WHERE em.key = 'chroma:document'
""")
# Then encode with BGE + NumPy cosine similarity for searchVector databases essentially store SQLite + float arrays; bypassing the API is viable. This highlights the flexibility of embedded architectures (ChromaDB, Milvus) vs. standalone services.
3. Distance vs. Similarity
ChromaDB returns distance (smaller = more similar), not similarity (larger = more similar). The code converts: (1 - distance) * 100 to produce a 0–100 relevance score for the frontend. Without conversion, higher relevance shows lower scores, confusing users.
4. Chunking Strategy Sets the Ceiling
Retrieval quality is bounded by chunk quality, not the embedding model. For legal documents with clear hierarchy, the project chunks by structure: first split by "Chapter X", then by "Article X", target ~500 characters per chunk with 50‑character overlap. This structural chunking outperforms fixed‑size sliding windows. Implementation in env_agent/init_chromadb.py: regex split chapters → split articles → merge short paragraphs to ~500 chars → pass 50‑char overlap to next chunk.
5. At Scale, Metadata Filtering Precedes Vector Search
For million‑scale data, filter by metadata (time, source, category) first to shrink the candidate set, then run vector search. Reversing the order hurts both speed and accuracy. Milvus and Qdrant support this hybrid filter natively. The effect is negligible in prototypes but must be designed into the architecture early.
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
