Do You Need a Vector Database for RAG? FastAPI Docs Benchmark Shows When Keyword Search Suffices
This article benchmarks keyword (BM25), vector (two Chinese embedding models), and hybrid retrieval on FastAPI's Chinese documentation, finding keyword search excels for exact-name queries while a larger vector model catches paraphrased questions, and hybrid helps only with strong embeddings; vector databases are unnecessary until data exceeds in-memory capacity.
Three Signals: What the Original Sources Actually Say
turbopuffer: RIP to the Primary Index, Not Vector Search
turbopuffer, a vector search cloud service used by Cursor and Notion, published a blog titled "RIP, vector database" on September 30. The article describes their storage architecture redesign: previously all data was stored according to the approximate nearest neighbor (ANN) index position, with attribute filtering and full-text search indexes attached to that position. This made vector search fast but hurt other queries — documents with multiple vectors required duplicate storage, vector re-clustering forced moving entire documents along with other indexes, and aggregation queries could only process small batches. In version 3 they switched to a new primary index, making ANN "just another" secondary index. Vector search remains, and they note a single index now scales to over 100 billion vectors. The "death" is the design that used vectors as the primary key, not vector search itself.
Claude Code: grep Wins for Code Retrieval
From a March 4 interview with Claude Code lead Boris Cherny on The Pragmatic Engineer: their "agentic search" is essentially glob and grep. The team tried local vector databases and recursive model-based indexing, but encountered issues like index staleness and complex permission handling. Model-driven glob + grep outperformed all alternatives. Two premises: the scenario is code, where function names, class names, and error messages are exact strings that grep matches perfectly; and grep is driven by the model, which iteratively reformulates keywords, examines results, and searches again, using multi-step reasoning to compensate for keyword search's inability to handle synonyms.
Karpathy: A Table of Contents Suffices for a Few Hundred Pages
Karpathy's April 4 GitHub project llm-wiki maintains interlinked Markdown notes with an index.md as a table of contents. The model reads the index first, then navigates to specific pages. On vector search, he writes that this approach works surprisingly well at medium scale — about 100 sources and a few hundred pages — eliminating the entire vector-based RAG infrastructure. However, when notes grow, real search is needed; the recommended tool qmd uses BM25 plus vector hybrid retrieval with model reranking. Together, the three signals show no claim that vector search is useless: turbopuffer changed its storage design, Claude Code addressed code retrieval, Karpathy addressed small-scale notes. Whether to use vector search depends on the problem shape; whether to deploy a vector database depends on data volume.
Retrieval Approaches
The retrieval step is where RAG most often fails: if the wrong documents are retrieved, the model's answer will be wrong regardless of quality. Common approaches:
Keyword search . The most common algorithm is BM25, scoring by term frequency and rarity in documents. Chinese requires word segmentation first, otherwise the whole sentence is treated as a single token. grep is also keyword search but without scoring, only presence/absence.
Vector search . An embedding model converts each text segment into a vector; semantically similar texts have nearby vectors. The query is also embedded, and the nearest vectors are retrieved. Advantage: recognizes rephrased queries. Cost: must run the model, store vectors, and recompute when documents change.
Hybrid search . Both methods run independently, then rankings are merged. A common fusion method is Reciprocal Rank Fusion (RRF): each document gets a score of 1/(60 + rank) from each list, summed and re-sorted.
Important distinction: vector search is a query method; a vector database is a dedicated service for storing vectors and performing approximate nearest neighbor search. Vector search does not require a vector database — when vector counts are low, they can reside in memory and be compared linearly. This benchmark uses exactly that approach, with no vector database.
Benchmark: FastAPI Official Chinese Documentation
Methodology
Corpus: 125 Markdown files from the FastAPI repository's Chinese docs, minus one translation test file and one banner, leaving 123 documents. Split by level-2 and level-3 headings; long segments further split at 400 characters. Total 1,577 segments.
Questions: 28 authored questions in two categories:
14 exact-name questions : contain API or parameter names directly, e.g., "How to use dependency_overrides", "WSGIMiddleware usage".
14 paraphrased questions : describe needs without API names, e.g., "User wants to upload an image to the backend" (correct answer: file upload page).
Ground truth: each question mapped to one or more correct documents in advance.
Three methods compared:
BM25: using jieba for Chinese tokenization, scoring per segment.
Vector: two local Chinese embedding models — smaller BAAI/bge-small-zh-v1.5 and larger jinaai/jina-embeddings-v2-base-zh — both run locally, no paid APIs.
Hybrid: BM25 and vector rankings merged via RRF.
All methods score segments, take the highest-scoring segment per document as the document score, then observe the rank of the correct document.
Results
BM25 is a surprisingly strong baseline. On exact-name questions it almost always ranks the correct document first — these questions are essentially looking for a specific term, which keyword search excels at.
The smaller bge-small model actually hurt performance: fewer exact-name questions ranked first compared to BM25. Switching to the larger jina model closed the gap on exact-name questions and placed all paraphrased questions in the top 5, while BM25 missed three.
Hybrid results depend on the paired model. Combined with bge-small, hybrid ranked fewer questions first than BM25 alone, dragged down by the weak model. Combined with jina, hybrid matched the best single method across the board.
With only 14 questions per category, a difference of one or two questions is not statistically decisive, but the trend is clear: keyword search dominates exact-name queries; a better vector model covers paraphrased queries.
Per-question illustration:
"WSGIMiddleware usage": BM25 and jina rank 1; bge-small ranks the correct document at 80, likely because the small model struggles with concatenated camel-case English identifiers.
"User wants to upload an image to the backend": BM25 ranks 36; both vector models rank 2. The document discusses file upload and UploadFile but never contains the words "image" or "backend", so keyword matching fails.
"Don't want outsiders to see API docs in production": only jina ranks 1. The document talks about disabling OpenAPI and the docs UI, sharing almost no words with the question.
An often overlooked observation: when writing paraphrased questions I only avoided API names, yet BM25 still ranked over half of them first. For example, "Return result to user first, then send email notification slowly" — the document's example happens to be about sending email. Real user questions often contain domain-specific terminology, so keyword search catches more than expected.
Cost of Vector Search
Vector search adds cost in three areas:
Index build time . BM25 builds in under 1 second; vector search must pass every segment through the model. The small model takes under a minute; the large model takes over 5 minutes. More segments means more computation.
Re-computation on updates . Changing one document requires re-embedding all its segments. With the large model, the longest document takes over ten seconds. If re-indexing is missed, stale content is retrieved — exactly the index staleness problem Claude Code mentioned.
Extra query step . Every query must first be embedded. The large model is an order of magnitude slower than the small one, but both are in the millisecond range.
These numbers are from a corpus of ~100 documents. At this scale all vectors fit in a few megabytes in memory; each query embeds the question and compares against 1,577 segments in ~10 ms, no vector database needed. At tens of thousands of documents, keyword search typically moves to a dedicated search engine, and vectors become too many for linear scan — that's when a vector database becomes necessary, a different scale not tested here.
Hybrid Search Core in a Few Lines
If adopting hybrid search, the ranking merge is simple. Below is the RRF implementation used in this test, taking multiple pre-sorted document lists:
def rrf(*ranks, k=60):
# Each document gets 1/(k+rank) from each list, summed and re-sorted
scores = {}
for rank in ranks:
for i, doc in enumerate(rank):
scores[doc] = scores.get(doc, 0) + 1 / (k + i + 1)
return sorted(scores, key=scores.get, reverse=True)The value k=60 comes from the original RRF paper. Larger k reduces the score gap between rank 1 and rank 10, preventing a single list's top result from overwhelming the other. Because RRF uses only ranks, not raw scores, BM25 scores and vector similarities need not be normalized to a common scale.
The results also remind us: hybrid is not always better. Mixing in a weak model can pull down questions that BM25 originally got right. Test each method individually first, then decide whether to combine.
How to Choose
Based on this test and the three original sources, my guidelines:
If the corpus fits entirely in the context window, skip retrieval and stuff it all in. (Personal experience, not tested here.) If it doesn't fit but is only a few hundred pages, follow Karpathy: write a table of contents and let the model read the index first, then drill down.
For code, config, or error lookups where the query contains exact identifiers, keyword search or grep is sufficient; let the model drive multiple search rounds.
When users ask in their own words and the documentation uses different terminology — e.g., customer support, internal knowledge bases — adding vector search is worthwhile; whether to also mix keyword search should be decided by testing.
Choose the embedding model carefully. In this test the gap between the two Chinese models was larger than the gap between vector and keyword search. Before production, test with your real questions rather than relying on leaderboards.
Concrete Product Options
The latter steps in the flowchart map to actual products: what to use for keyword search at scale, where to put vectors when they exceed memory, which embedding model to pick. The three tables below are compiled from official documentation and repositories; none were benchmarked in this article. Only jieba+BM25 and the two local embedding models were actually tested.
First, vector databases — for when vectors outgrow in-memory linear scan:
Second, keyword and hybrid search engines — for when documents exceed a few tens of thousands and keyword search needs a dedicated engine:
Third, embedding models — for the "pick embedding model and test again" step:
Several observations from the tables:
The line between vector databases and search engines is blurring. Elasticsearch supports kNN on vectors, Milvus natively supports BM25, Qdrant uses sparse vectors for exact term matching, and pgvector can pair with Postgres full-text search. If you already run Postgres or Elasticsearch, you can add vectors there and avoid maintaining a separate service.
Chinese tokenization must be verified separately. BM25 quality depends on segmentation; this test used jieba. Elasticsearch's built-in CJK analyzer only splits Chinese into bigrams; word-level segmentation requires a plugin. Meilisearch's built-in Chinese tokenizer is based on jieba. Postgres's built-in full-text search has no Chinese configuration.
Switching embedding models requires re-embedding all segments. The jina model took over 5 minutes to index this corpus; larger corpora will take proportionally longer. None of the models in the table were tested on these questions — suitability must be validated with your own real queries.
Returning to the title question: Does RAG still need a vector database? Two-layer answer. Vector search is useful — it covers paraphrased questions — but it should not be the first step. Start with keyword search as a baseline, test with real questions to see what it misses, then decide whether to add vector search. A vector database is only needed when the data volume exceeds what can be compared in memory; at the scale of a hundred-odd documents like this test, keeping vectors in memory is sufficient.
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
Tech Ocean
Focused on AI programming, sharing ready-to-use development efficiency solutions.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
