Databases 20 min read

Dameng DM8 + Easysearch Full-Text Search Validation: 12x Speedup on 1M Chinese Docs

This report validates a read-write separation architecture using Dameng DM8 for transactions and Easysearch for full-text search, syncing 1 million Chinese documents via Logstash; full sync completes in 5.5 minutes, incremental changes propagate within 30 seconds, and search latency drops 12x compared to SQL LIKE while maintaining identical recall.

Mingyi World Elasticsearch
Mingyi World Elasticsearch
Mingyi World Elasticsearch
Dameng DM8 + Easysearch Full-Text Search Validation: 12x Speedup on 1M Chinese Docs

Conclusion Summary

Architecture feasible, end-to-end pipeline works: Dameng DM8 → Logstash → Easysearch link fully operational.

Full sync performance meets target: 1,000,000 Chinese documents synchronized in ~5.5 minutes (~3,000 docs/sec).

Incremental sync latency acceptable: Inserts, updates, and logical deletes visible within 30 seconds (polling interval 1 minute).

Search performance vastly superior: Dameng LIKE average 1,322 ms vs Easysearch average 107 ms — ~12x speedup.

Speedup without recall loss: Five equivalent semantic queries return identical hit counts as Dameng LIKE.

One-sentence summary: The four-step architecture is completely correct and has been implemented step by step. Dameng retains transactional duties; Easysearch adds search capability; business systems require no changes. Before external delivery, upgrade sync layer from polling to CDC and replace sync tooling with domestic alternatives.

Background and Goals

Dameng DM8 is a mainstream domestic relational database widely used in government, finance, and energy sectors. However, relational databases have inherent limitations in full-text search scenarios, struggling with complex retrieval needs over unstructured text.

Read-write separation approach: Dameng remains the transactional store and source of truth. Text data is streamed in real time to Easysearch via a synchronization pipeline, adding high-performance full-text search without altering the existing business architecture.

Validation broken into four steps:

Download and install Dameng DM8 locally.

Download and install Logstash locally.

Establish Logstash-to-Easysearch data sync.

Implement full-text search and quantify comparison.

Test Environment

Dameng Database: DM8 (Server 64 V8), charset GB18030, deployed at C:\dmdbms, port 5236.

Easysearch: 2.4.0 (kernel Elasticsearch 7.10.2, Lucene 9.12.2), endpoint http://127.0.0.1:9200.

Logstash: 9.0.0 (output plugin 12.0.2), path D:\software\logstash-9.0.0-windows-x86_64.

Dameng JDBC Driver: DmJdbcDriver8.jar placed in logstash/vendor/jar/jdbc/.

Chinese Tokenization: analysis-ik + analysis-pinyin as Easysearch plugins.

OS: Windows (single-node full-stack deployment).

Test Data: KB_ARTICLE table, 1,000,000 Chinese documents with fields: title, content, category, author, keywords, updated_at, is_deleted; date range 2024-09-29 to 2026-09-28; ~10,000 logically deleted rows.

Architecture

Architecture diagram
Architecture diagram

Core principle: Writes still go to Dameng , search queries route to Easysearch, zero changes to original database, business unaware .

Verification Process and Results

4.1 Dameng DM8 Deployment and Data Preparation

DM8 installed and instance initialized; service DmServiceDMSERVER running normally.

Charset set to GB18030 ; Chinese read/write works (note: SQL scripts must be submitted to DIsql in GBK encoding to avoid garbled characters).

Bulk data generation using CONNECT BY LEVEL; 1 million rows loaded in ~40 seconds.

Data preparation screenshot
Data preparation screenshot

4.2 Sync Link Establishment

This was the highest technical hurdle in the entire validation.

Easysearch reports version.number = 2.4.0 by default, but Logstash's elasticsearch output plugin requires Elasticsearch ≥ 7.0.0, causing immediate startup failure:

LogStash::ConfigurationError:
Could not connect to a compatible version of Elasticsearch

Switching Logstash versions does not solve it — plugin generations 9/10/11/12 require minimum Elasticsearch versions 6.6.0 / 6.6.0 / 7.0.0 / 7.0.0 respectively, all higher than 2.4.0.

Solution: Enable Easysearch's API compatibility mode in config/easysearch.yml:

elasticsearch.api_compatibility: true
elasticsearch.api_compatibility_version: "7.10.2"
Compatibility mode config
Compatibility mode config

After restart, Easysearch reports standard Elasticsearch version and handshake succeeds. This config does not support dynamic reload; must edit file and restart node.

4.3 Full Sync

Documents: 1,000,000
Duration: ~5.5 minutes
Throughput: ~3,000 docs/sec
Index size: 167 MB (zstd compression)
Concurrency config: -w 4 -b 2000

4.4 Incremental Sync

On Dameng side executed "insert 3 rows + update 1 row + logical delete 1 row"; all synced to Easysearch within 30 seconds (scheduling cycle 1 minute).

Incremental sync verification
Incremental sync verification

Insert: ✅ 3 new records searchable.

Update: ✅ Title and keywords updated in sync.

Logical Delete: ✅ is_deleted set to 1, search side filters via is_deleted = 0.

Important note: Logstash JDBC input uses timestamp-based polling, cannot capture physical DELETE . Therefore business tables must include is_deleted + updated_at logical delete fields; on delete, rewrite both fields; search side filters with is_deleted = 0 .

4.5 Full-Text Search Capabilities

Using "中文分词" as query term, index hits 166,667 documents; highlighting works (example):

id=994253 [国产化] 索引优化技术方案说明(编号 DM-994253)
  <em>中文</em><em>分词</em>器的词典规模会直接影响检索的准确率与内存占用。
  关锥词高亮功能显营改善了用户在知识库中的检索体验。

Verified search capabilities:

Keyword search + relevance scoring (BM25)

Keyword highlighting

Multi-field weighted search (title weight ×2)

Multi-condition combination (must / filter / must_not)

Search-as-aggregation (single query returns category distribution, top-N authors)

4.6 Performance Comparison

Comparison scope: Dameng CONTENT LIKE '%keyword%' vs Easysearch match_phrase (semantically equivalent, both mean "contains this continuous string").

Performance comparison chart
Performance comparison chart
Query Phrase                     | Dameng LIKE (ms) | Hits | Easysearch (ms) | Hits
---------------------------------|------------------|------|-----------------|------
全文检索                         | 1,119            | 83,334 | 37              | 83,334
中文分词                         | 1,071            | 166,667| 35              | 166,667
国产搜索引擎在政务和金融领域     | 1,183            | 83,333 | 143             | 83,333
达梦数据库与 Easysearch 协同部署 | 2,002            | 83,333 | 233             | 83,333
向量检索为大模型问答提供         | 2,346            | 83,333 | 84              | 83,333
Average                          | 1,322            | —    | 107             | —

Conclusion: ~12x average speedup, and hit counts across all five queries exactly match Dameng LIKE — speedup not achieved by sacrificing recall.

Speedup summary
Speedup summary
Note: Dameng latency affected by OS cache; cold cache average ~1,897 ms (up to 17x speedup). Easysearch latency stable, single-query range 35–233 ms.

4.7 Capability Differences: What LIKE Cannot Do

Search Requirement                              | Dameng LIKE Hits | Easysearch Hits
------------------------------------------------|------------------|----------------
全文检索 中文刉词 (多关锥词组合)         | 0                | 666,669
达梦 Easysearch 协同部署 (越词序、隔词)    | 0                | 250,001
'%全文检索%协同部署%' (词序题倒)     | 0                | 83,333

Dameng LIKE requires literal continuous match (tokenized terms don't match) ; multi-keyword must be expressed as long OR + LIKE chains. Easysearch, after IK tokenization, searches by terms, natively supporting multi-word combos and order-independent matching. This is the fundamental difference between full-text search and string matching.

Key Issues and Resolutions

# | Issue                                      | Impact                     | Resolution
--|--------------------------------------------|----------------------------|------------
1 | Easysearch reports 2.4.0, Logstash refuses | Pipeline cannot start      | Enable api_compatibility & restart Easysearch
2 | Dameng JDBC driver class name format       | Driver load failure        | Must prefix with Java:: → Java::dm.jdbc.driver.DmDriver
3 | Logstash 9 missing vendor/jar/jdbc dir     | No place to put driver     | Manually create dir and place DmJdbcDriver8.jar
4 | Output xpack capabilities incompatible     | Startup error              | Set manage_template=false, ilm_enabled=false
5 | Polling cannot capture physical DELETE     | Dirty data in search       | Add is_deleted + updated_at, use logical delete
6 | Same-timestamp batch boundary loses data   | Incomplete sync            | Use >= instead of > in sync SQL, with document_id for idempotency
7 | IK index/search analyzer mismatch          | Long-phrase hits = 0       | Unify both to ik_max_word; close/open index to apply
8 | IK overlapping sub-terms cause gaps        | ~50% short-term miss       | Add "slop": 1 to match_phrase
Issues 7 and 8 are two independent defects ; must be handled separately. In testing, even with slop set to 5, hits remained 0 when analyzers mismatched — they cannot substitute for each other.

Capability Boundaries and Risks

6.1 Sync Timeliness

Logstash JDBC is minute-level polling , not "real-time sync". Easysearch officially recommends "Dameng DM8 + CDC real-time sync + Easysearch" for Dameng scenarios, achieving second-level precision with certified compatibility.

6.2 Component Domestication

Logstash is an Elastic product (SSPL/ELv2 license). In a "Dameng domestic + Easysearch domestic" full-stack Xinchuang (信创) solution, a non-domestic component creates a narrative gap. External materials should replace Logstash with INFINI Gateway or SeaTunnel ; keep Logstash only for internal validation.

6.3 Chinese Tokenization Dictionary

IK built-in dictionary lacks coverage for industry-specific terms (e.g., "达梦" splits into "达" and "梦" as two single characters).

IK tokenization example
IK tokenization example

Production environments should extend via config/analysis-ik/custom/*.dic with custom dictionaries.

Conclusions and Recommendations

7.1 Validation Conclusion

Original four-step skeleton completely correct , all items landed and passed:

Dameng DM8 local installation ✅ Done

Logstash local installation & config ✅ Done

Logstash → Easysearch data sync ✅ Full + incremental verified

Full-text search capability ✅ Search, highlight, aggregation, combined queries all work

Key performance metrics: 1M rows full sync 5.5 min, incremental visible within 30 sec, search 12–17x faster than LIKE with identical recall.

7.2 Next Steps: Two-Layer Approach

L1 · Minimum Viable Validation (already achieved) — Logstash + logical delete, quick demo for pipeline validation and mapping design freeze.

L2 · Production Alignment (recommended)

Replace sync layer with CDC (Dameng DMDRS/DMHS or SeaTunnel → Kafka → Easysearch) to align with official certified pipeline.

Add sync latency monitoring and checkpoint resume mechanism.

Complete domestic replacement of sync tooling to close Xinchuang narrative.

Appendix A: Key Configurations

Sync Pipeline Core Config (Logstash)

input {
  jdbc {
    jdbc_driver_library  => ".../vendor/jar/jdbc/DmJdbcDriver8.jar"
    jdbc_driver_class    => "Java::dm.jdbc.driver.DmDriver"
    jdbc_connection_string => "jdbc:dm://127.0.0.1:5236"
    jdbc_user            => "SYSDBA"
    statement            => "SELECT * FROM KB_ARTICLE WHERE UPDATED_AT >= :sql_last_value"
    use_column_value     => true
    tracking_column      => "updated_at"
    tracking_column_type => "timestamp"
    schedule             => "*/1 * * * *"
  }
}

output {
  elasticsearch {
    hosts         => ["http://127.0.0.1:9200"]
    index         => "kb_article"
    document_id   => "%{id}"
    manage_template => false
    ilm_enabled   => false
  }
}

Index Mapping Highlights

{
  "settings": {
    "number_of_shards": 1,
    "number_of_replicas": 0,
    "index.codec": "zstd",
    "index.source_reuse": true
  },
  "mappings": {
    "properties": {
      "title":      { "type": "text", "analyzer": "ik_max_word", "search_analyzer": "ik_max_word" },
      "content":    { "type": "text", "analyzer": "ik_max_word", "search_analyzer": "ik_max_word" },
      "category":   { "type": "keyword" },
      "is_deleted": { "type": "integer" },
      "updated_at": { "type": "date" }
    }
  }
}

Appendix B: Deliverables List

README.md

— Reproduction manual blog-dameng-easysearch.md — Chinese technical blog (includes troubleshooting) easysearch/dsl-06-ik-analyzer.md — Executable DSL collection for IK analyzer alignment easysearch/kb_article_mapping.json — Index mapping template logstash/dameng-to-easysearch.conf — Sync pipeline config tools/benchmark.py — Performance comparison benchmark script tools/dsl_verify.py — 21 DSL batch regression script sql/01_schema.sql — Table creation statements sql/load_template.sql — Test data generation template

Report based on local full-stack hands-on testing; all data reproducible.

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

performance benchmarkdata synchronizationCDCfull-text searchdomestic databaseLogstashChinese tokenizationIK analyzerEasysearchDameng DM8
Mingyi World Elasticsearch
Written by

Mingyi World Elasticsearch

The leading WeChat public account for Elasticsearch fundamentals, advanced topics, and hands‑on practice. Join us to dive deep into the ELK Stack (Elasticsearch, Logstash, Kibana, Beats).

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.