Easysearch 2.4.0 Vector Search: Verified Guide to Indexing, Querying & Hybrid Search Pitfalls
This article provides a step-by-step verified guide to implementing vector search in Easysearch 2.4.0, covering index creation with dense_vector fields, data ingestion with embeddings, k-NN query pitfalls including the mandatory 'k' parameter, and the critical distinction between compound queries and RRF-based hybrid search for combining keyword and semantic search.
Understanding Vector Search in Easysearch 2.4.0
Traditional search matches keywords exactly; vector search matches meaning. An embedding model (e.g., bge-m3) converts text into a dense vector (e.g., 1024 dimensions). Semantically similar texts produce vectors that are mathematically close. Easysearch 2.4.0 embeds this capability natively, removing the need for a separate k-NN plugin.
Step 1: Create an Index with a dense_vector Field
Define the index mapping with a dense_vector field. Key parameters: dims: Vector length, must match the embedding model output (e.g., 1024 for bge-m3). Mismatch causes write errors. similarity: Distance metric; cosine is standard for semantic search. Other options ( dot_product, l2_norm, max_inner_product) are rarely needed. index: true: Enables search on this field. Set false if only storing vectors.
PUT /product_kb
{
"mappings": {
"properties": {
"title": { "type": "text" },
"content": { "type": "text" },
"category": { "type": "keyword" },
"content_vector": {
"type": "dense_vector",
"dims": 1024,
"similarity": "cosine",
"index": true
}
}
}
}Step 2: Write Documents with Pre-computed Vectors
Easysearch does not generate vectors; you must compute them externally using an embedding model and include the vector array in the document:
POST /product_kb/_bulk
{ "index": { "_id": "1" } }
{ "title": "Easysearch Vector Search", "content": "Native HNSW supports semantic search", "category": "release", "content_vector": [0.021, -0.134, 0.087, ...] }The vector field is a plain float array; no special handling is required.
Step 3: Search — Two Syntaxes and a Critical Pitfall
Convert the user query to a vector (same embedding model) and run a k-NN search. Two syntaxes exist:
Syntax A: knn inside query (requires mandatory k )
POST /product_kb/_search
{
"query": {
"knn": {
"field": "content_vector",
"query_vector": [0.02, -0.13, 0.08, ...],
"k": 5,
"num_candidates": 50
}
},
"size": 5
}Pitfall: Omitting k causes a misleading error: "[knn] queries are only supported on [dense_vector] fields". In Easysearch 2.4.0, k is mandatory for this syntax; the parser fails before validating the field type. k = final results wanted; num_candidates = per-shard candidate pool (typically 5–10× k).
Syntax B: knn at top level (recommended)
POST /product_kb/_search
{
"knn": {
"field": "content_vector",
"query_vector": [0.02, -0.13, 0.08, ...],
"k": 5,
"num_candidates": 50
},
"size": 5
}This syntax is cleaner and avoids the k parsing issue.
Step 4: Combining Keyword and Vector Search — Compound Query vs. Hybrid Search
Real-world search often needs both keyword (BM25) and semantic (vector) matching. Two distinct approaches exist:
Compound Query (bool + knn + match)
Place both queries in a bool clause; scores are summed directly.
POST /product_kb/_search
{
"size": 10,
"query": {
"bool": {
"must": [
{ "knn": { "field": "content_vector", "query_vector": [...], "k": 10, "num_candidates": 100 } }
],
"should": [
{ "match": { "content": "vector search" } }
]
}
}
}Problem: BM25 scores and cosine similarity scores are on different scales; adding them directly lets one dominate. You must manually tune boost parameters — a trial-and-error process.
Hybrid Search (RRF via search pipeline)
Official term for a more robust method: each sub-query produces a ranked list; Reciprocal Rank Fusion (RRF) merges the ranks, eliminating scale mismatch.
1. Create a search pipeline with an RRF ranker:
PUT /_search/pipeline/rrf-pipeline
{
"rerank_processors": [
{ "hybrid_ranker_processor": { "combination": { "technique": "rrf", "rank_constant": 60 } } }
]
}2. Execute a hybrid query referencing the pipeline:
GET /product_kb/_search?search_pipeline=rrf-pipeline
{
"query": {
"hybrid": {
"queries": [
{ "match": { "content": "vector search" } },
{ "knn": { "field": "content_vector", "query_vector": [...], "k": 10, "num_candidates": 100 } }
]
}
}
}RRF merges rankings without score normalization; rank_constant (default 60) controls how much lower-ranked items can influence the fused result. This approach avoids weight tuning and is recommended when ranking quality matters.
Key Takeaways
Align dims with your embedding model (e.g., 1024 for bge-m3).
Always include k in knn queries (especially inside query).
Choose compound query for simplicity with manual boosting, or hybrid search (RRF) for parameter-free, scale-invariant fusion.
All code examples have been executed and verified in the Easysearch console; adapt field names and vector dimensions to your schema.
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
Mingyi World Elasticsearch
The leading WeChat public account for Elasticsearch fundamentals, advanced topics, and hands‑on practice. Join us to dive deep into the ELK Stack (Elasticsearch, Logstash, Kibana, Beats).
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
