Databases 16 min read

Dameng DM8 + Easysearch: 17x Faster Full-Text Search via Logstash Sync

This article details integrating Dameng DM8 database with Easysearch 2.4.0 via Logstash 9.0 for full-text search, covering version compatibility fixes, JDBC driver configuration, logical deletion handling, timestamp boundary issues, IK analyzer alignment for Chinese text, and benchmark results showing 17x faster queries with consistent recall versus LIKE.

Mingyi World Elasticsearch
Mingyi World Elasticsearch
Mingyi World Elasticsearch
Dameng DM8 + Easysearch: 17x Faster Full-Text Search via Logstash Sync

Why Add a Search Engine to Dameng Database

Relational databases excel at transactions, not retrieval. Using LIKE '%keyword%' for full-text search triggers a full table scan plus string matching, suffering from three fatal flaws: it is slow (scans entire table regardless of query), inaccurate (requires literal contiguous match, failing on Chinese word segmentation), and unranked (returns hits in primary-key order without relevance scoring). At million-row scale, all three problems explode simultaneously.

The solution is read-write separation: Dameng remains the transactional source of truth, while text data is synchronized to Easysearch for indexing, tokenization, and scoring. The business system architecture stays unchanged, no extra database indexes are needed, and search capability upgrades independently.

Architecture: Dameng DM8 (transaction/source) → Sync pipeline → Easysearch (search/tokenization/scoring)

Environment & Four-Step Goal

Dameng Database : DM8, charset GB18030, location C:\dmdbms, port 5236

Logstash : 9.0.0, location D:\software\logstash-9.0.0-windows-x86_64 Easysearch : 2.4.0 (kernel ES 7.10.2), endpoint http://127.0.0.1:9200 Chinese Tokenizer : analysis-ik (Easysearch plugin)

Goal: install Dameng, install Logstash, establish sync, implement full-text search. The real time is spent in pitfalls, not installation.

Pitfall 1: Easysearch Version Number Rejects Logstash Handshake

Symptom: Pipeline syntax validates but startup fails with

LogStash::ConfigurationError: Could not connect to a compatible version of Elasticsearch

.

Cause: The logstash-output-elasticsearch plugin performs a version handshake on startup. Easysearch reports version 2.4.0, but the plugin requires Elasticsearch ≥ 7.0.0. To the plugin, 2.4.0 looks like a 2014 antique and is rejected.

Wrong Direction: Changing Logstash version does not help. Plugin versions 9/10/11/12 require minimum Elasticsearch 6.6.0/6.6.0/7.0.0/7.0.0 — all greater than 2.4.0. Downgrading to Logstash 7.x still fails.

Solution: Easysearch provides an API compatibility switch to masquerade as standard Elasticsearch. Edit config/easysearch.yml:

elasticsearch.api_compatibility: true
elasticsearch.api_compatibility_version: "7.10.2"

After restart, curl http://127.0.0.1:9200/ returns version 7.10.2 (build_flavor: default, lucene 8.7.0) and handshake succeeds immediately.

Three Reminders:

This config does not support dynamic modification. Using _cluster/settings API throws an error; must edit file and restart.

Backup before change, then compare version response before/after restart to confirm.

This alters the existing environment; document it clearly so teammates aren't confused by the "mysterious 7.10.2".

Four Critical Details Connecting Dameng to Logstash

Core pipeline configuration:

input {
  jdbc {
    jdbc_driver_library => "D:/software/logstash-9.0.0-windows-x86_64/logstash-9.0.0/vendor/jar/jdbc/DmJdbcDriver8.jar"
    jdbc_driver_class => "Java::dm.jdbc.driver.DmDriver"
    jdbc_connection_string => "jdbc:dm://127.0.0.1:5236"
    jdbc_user => "SYSDBA"
    jdbc_password => "******"
    statement => "SELECT * FROM KB_ARTICLE WHERE UPDATED_AT >= :sql_last_value"
    use_column_value => true
    tracking_column => "updated_at"
    tracking_column_type => "timestamp"
    schedule => "*/* * * * *"
  }
}
output {
  elasticsearch {
    hosts => ["http://127.0.0.1:9200"]
    index => "kb_article"
    document_id => "%{id}"
    manage_template => false
    ilm_enabled => false
  }
}

Four Must-Watch Points:

① Driver class name must carry Java:: prefix. Writing dm.jdbc.driver.DmDriver causes class-not-found; correct is Java::dm.jdbc.driver.DmDriver.

② Logstash 9 lacks vendor/jar/jdbc/ directory. Old tutorials tell you to drop the driver there, but the directory doesn't exist in new versions — create it manually:

mkdir -p vendor/jar/jdbc
# then copy DmJdbcDriver8.jar into it

③ Connection string is jdbc:dm://host:5236 . Port 5236, default user SYSDBA. DM8 enforces password complexity policy ( PWD_POLICY); relax it in PoC environments or database creation will hang.

④ Output side: two switches must be off. manage_template => false (prevent Logstash from modifying Easysearch index templates) and ilm_enabled => false. Leaving them on causes startup errors about unrecognized settings.

Two Silent Pitfalls in Sync Logic

These produce no errors but silently lose data or leave dirty data — the kind a customer can expose with a single question.

Pitfall A: Polling Misses DELETE Operations

JDBC input uses tracking_column (timestamp) for incremental pulls; it only sees "new rows" and "updated rows". If someone deletes a row in Dameng, Logstash never knows — the record disappears from the database but remains in Easysearch. A customer searches and finds already-deleted data; trust evaporates instantly.

Fix: Logical Deletion. Add two columns to source table: is_deleted: set to 1 on delete, no physical removal. updated_at: must also update timestamp on delete, otherwise sync never detects the change.

Sync SQL includes the time condition; query-time filtering on business side:

{
  "filter": [ { "term": { "is_deleted": 0 } } ]
}

Physical deletion is CDC's job; polling cannot achieve it.

Pitfall B: Timestamp Boundary Must Use >= Not >

Using > on tracking_column seems logical, but 1 million rows only have ~730 distinct timestamps (second precision means many rows share the same second). If a batch stops mid-timestamp, rows with that timestamp not yet synced are permanently skipped.

Using >= lets same-timestamp rows be fetched again. Duplicate writes are harmless because output uses id for upsert; missing writes are fatal.

Chinese Tokenization: IK's Two Analyzers Must Align

Symptom: Both index and search use IK, but long-phrase queries miss massively — searching "domestic search engine in government and finance" returns 0 hits, while Dameng LIKE returns 80,000+.

Cause: ik_max_word performs fine-grained segmentation, producing overlapping sub-tokens. Example:

"Chinese tokenizer dictionary scale" → Chinese(0), tokenizer(1), token(2), izer(3) ... ← overlapping positions
Query "Chinese token" → Chinese(0), token(1) ← two tokens separated by 2 positions
match_phrase

requires tokens to be position-contiguous in the document. ik_max_word 's overlapping tokens scramble positions, so short query phrases fail to match.

Two Fixes, Both Applied for Safety:

Unify index and search analyzers. Default mapping often uses analyzer: ik_max_word + search_analyzer: ik_smart; this combo clashes on phrase queries. Standardize on ik_max_word for both.

Add slop to match_phrase to allow slight position offsets:

{
  "match_phrase": {
    "content": {
      "query": "Chinese token",
      "slop": 1
    }
  }
}

After fix, hit counts match Dameng LIKE exactly. Changing search_analyzer is a mapping change; hot reload doesn't work due to mapping cache. Force reload via close/open index (no data rebuild):

curl -XPOST "http://127.0.0.1:9200/kb_article/_close"
curl -XPOST "http://127.0.0.1:9200/kb_article/_open"

Real-World Test Results

Full Sync : 1,000,000 rows, 5.5 minutes (≈3,000 docs/s)

Incremental Sync : Insert/update/logical delete visible within 30 seconds

Index Size : 1,000,003 docs, 166.7 MB

Search Performance : Dameng LIKE avg 1897ms → Easysearch avg 110ms , ~17x faster

Recall Consistency : Five equivalent semantic queries, hit counts identical to LIKE

Crucially, speedup did not sacrifice recall. Easysearch also handles what LIKE cannot: multi-keyword combination "full-text search Chinese tokenization" — Dameng LIKE returns 0 hits (requires contiguous words), Easysearch returns 666,667 hits . This is the fundamental difference between full-text search and string matching.

Capability Boundaries of This Solution

Two facts must be stated upfront to avoid awkward questions:

First, Logstash polling is minute-level; cannot claim "real-time sync". Official recommended path for Dameng is DM8 + CDC real-time sync + Easysearch (second-level precision, already compatibility-certified). Dameng side can use DMDRS/DMHS or third-party SeaTunnel. Logstash JDBC is a "Day-1 Demo connector", not a end-state solution.

Second, Logstash is not a domestic component. Dameng is domestic, Easysearch is domestic, but the middle layer is Elastic's product (SSPL/ELv2 license) — a visible gap in full-stack xinchuang (indigenous innovation) narratives. External materials should replace Logstash with INFINI Gateway or SeaTunnel; keep Logstash only for internal validation.

Summary

The real difficulty wasn't "can we connect" but four silent failure points: version masquerading, delete semantics, timestamp boundaries, IK analyzer alignment. Once those are cleared, the rest is grunt work.

A pragmatic rollout rhythm:

L1 Minimum Validation: Logstash + logical deletion, same-day demonstrable search results, lock down pipeline and mapping.

L2 Production Hardening: Swap sync layer to CDC, add sync latency monitoring and checkpoint resume, remove Logstash from architecture diagram.

Why bother? 1897ms vs 110ms, and multi-keyword from 0 hits to 660k hits — those two numbers speak for themselves.

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

JDBCdata synchronizationversion compatibilityfull-text searchLogstashChinese tokenizationIK analyzerDamengDM8Easysearch
Mingyi World Elasticsearch
Written by

Mingyi World Elasticsearch

The leading WeChat public account for Elasticsearch fundamentals, advanced topics, and hands‑on practice. Join us to dive deep into the ELK Stack (Elasticsearch, Logstash, Kibana, Beats).

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.