Databases 12 min read

Beyond RAG: Vector Databases as General-Purpose Similarity Search Engines

This article explains why vector databases are general-purpose similarity search engines beyond RAG, covering embedding model selection (BGE), non-RAG use cases like recommendation and anomaly detection, core operations with ChromaDB and HNSW, selection criteria for ChromaDB, FAISS, Milvus, and Pinecone, and five practical engineering lessons on recall evaluation, API workarounds, distance vs similarity, chunking strategies, and metadata filtering.

Linyb Geek Road
Linyb Geek Road
Linyb Geek Road
Beyond RAG: Vector Databases as General-Purpose Similarity Search Engines

Why Vector Databases Are Needed

Traditional keyword search (Elasticsearch, SQL LIKE) matches literal text. A query for "印染废水 COD 限值" only finds documents containing those exact characters, missing synonyms, phrasing variations, or semantic equivalents. Vector search matches semantics: both the query and a document like "纺织染整工业水污染物排放标准规定 COD 排放限值 200mg/L" are embedded into high‑dimensional vectors, and cosine similarity retrieves them even with zero shared terms.

The challenge is scale: exact nearest‑neighbor search is O(Nd) linear scan, feasible for 10k vectors but not 100M. Vector databases solve this with ANN (Approximate Nearest Neighbor) indexes — sacrificing a tiny amount of precision for sub‑linear query time, returning results that are "close enough" rather than mathematically exact.

Embedding: Turning Text into Vectors

Vector databases store float arrays, not text. Embedding models compress discrete text into a continuous high‑dimensional space where semantically similar texts are neighbors. The project uses BGE (BAAI/bge-small-zh-v1.5) for three reasons:

Chinese performance: trained specifically for Chinese, not an English model with added Chinese support.

Small footprint: ~93 MB, runs on CPU without GPU.

Industry adoption: mature community, strong reproducibility.

A critical detail: BGE requires different prefixes for queries and documents. The query side adds an instruction prefix:

emb = self._model.encode(
    f"为这个句子生成表示以用于检索相关文章:{text}"
)

Documents are encoded without the prefix. Forgetting the prefix drops recall by 5–15 percentage points while the code still runs, making the regression hard to detect. This also illustrates that embedding models differ in how they define similarity — BGE separates query intent from document content, others do not.

Beyond RAG: Other Vector Database Use Cases

Viewing vector databases as general similarity search engines reveals many non‑RAG applications:

Vector database non‑RAG use cases: recommendation, deduplication, anomaly detection, clustering visualization, long‑context compression
Vector database non‑RAG use cases: recommendation, deduplication, anomaly detection, clustering visualization, long‑context compression

Recommendation systems: user history embedding + item embedding → nearest neighbors. Powers "you may like", "similar videos", "related news" without LLMs or complex rule engines.

Document deduplication & anomaly detection: two sides of the same problem. Embed known normal samples; new samples with similarity >0.95 are duplicates (news crawlers skip them), similarity <0.3 are anomalies (industrial IoT flags abnormal logs). Same formula, different thresholds.

Long‑context compression: split a large report or contract into N chunks, embed all, retrieve only the top‑3 relevant chunks for the LLM. The vector database acts as a selector for limited context windows.

Abstracting the vector database as team infrastructure yields more reuse than a one‑line import in a RAG tutorial.

Basic Vector Database Operations

APIs across vector databases share a common skeleton: create index, query, delete. ChromaDB example:

# Create
client = chromadb.PersistentClient(path="./data/chroma_db")
collection = client.get_or_create_collection("env_laws")

# Insert
collection.add(documents=chunks, embeddings=embeddings, metadatas=metadatas, ids=ids)

# Query
results = collection.query(query_embeddings=[query_vec], n_results=5)

The core indexing algorithm is HNSW (Hierarchical Navigable Small World) . Analogy: a city transport network — sparse express highways (top layer) quickly bring you to the target district, then dense local streets (bottom layer) find the exact neighbor. Query starts at the top layer, complexity approaches O(log N). Milvus, Qdrant, Weaviate all default to HNSW or variants.

However, selection hinges on engineering traits (deployment model, scalability, ecosystem), not which index algorithm is newest.

Mainstream Vector Database Selection

Vector database landscape: embedded (ChromaDB, FAISS, numpy) vs service (Milvus, Weaviate, Qdrant, Pinecone)
Vector database landscape: embedded (ChromaDB, FAISS, numpy) vs service (Milvus, Weaviate, Qdrant, Pinecone)

Practical guidance by scale:

Tens to hundreds of thousands → ChromaDB to get running fast.

Performance‑sensitive, single‑node → FAISS .

Billion‑scale, distributed → Milvus .

Zero‑ops preference → Pinecone (managed service).

The landscape splits into embedded (ChromaDB, FAISS, raw numpy) and service‑oriented (Milvus, Weaviate, Qdrant, Pinecone) architectures.

Five Practical Lessons

1. Evaluation Must Go Beyond "It Runs"

Before launch, measure recall: the fraction of truly relevant documents retrieved. Method: manually annotate 50–100 "question → relevant document" pairs, run retrieval, compute hit rate.

2. Don’t Wrestle with Broken APIs

ChromaDB 0.5.5 clashed with huggingface-hub versions, causing collection.query() errors. Instead of debugging, the team bypassed ChromaDB’s query API: read documents directly from the underlying chroma.sqlite3, then used sentence-transformers + NumPy cosine similarity for retrieval:

# Read directly from ChromaDB SQLite, skip query API
cursor.execute("""
    SELECT em.string_value, em2.string_value, em3.string_value
    FROM embedding_metadata em
    LEFT JOIN embedding_metadata em2 ON em.id = em2.id AND em2.key = 'law_name'
    LEFT JOIN embedding_metadata em3 ON em.id = em3.id AND em3.key = 'source'
    WHERE em.key = 'chroma:document'
""")
# Then encode with BGE + NumPy cosine similarity for search

Vector databases essentially store SQLite + float arrays; bypassing the API is viable. This highlights the flexibility of embedded architectures (ChromaDB, Milvus) vs. standalone services.

3. Distance vs. Similarity

ChromaDB returns distance (smaller = more similar), not similarity (larger = more similar). The code converts: (1 - distance) * 100 to produce a 0–100 relevance score for the frontend. Without conversion, higher relevance shows lower scores, confusing users.

4. Chunking Strategy Sets the Ceiling

Retrieval quality is bounded by chunk quality, not the embedding model. For legal documents with clear hierarchy, the project chunks by structure: first split by "Chapter X", then by "Article X", target ~500 characters per chunk with 50‑character overlap. This structural chunking outperforms fixed‑size sliding windows. Implementation in env_agent/init_chromadb.py: regex split chapters → split articles → merge short paragraphs to ~500 chars → pass 50‑char overlap to next chunk.

5. At Scale, Metadata Filtering Precedes Vector Search

For million‑scale data, filter by metadata (time, source, category) first to shrink the candidate set, then run vector search. Reversing the order hurts both speed and accuracy. Milvus and Qdrant support this hybrid filter natively. The effect is negligible in prototypes but must be designed into the architecture early.

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

RAGvector-databaseMilvusHNSWembeddingsimilarity-searchChromaDBchunking strategy
Linyb Geek Road
Written by

Linyb Geek Road

Tech notes

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.