Tagged articles

Chinese tokenization

5 articles · Page 1 of 1
Mingyi World Elasticsearch
Mingyi World Elasticsearch
Sep 28, 2026 · Databases

Dameng DM8 + Easysearch Full-Text Search Validation: 12x Speedup on 1M Chinese Docs

This report validates a read-write separation architecture using Dameng DM8 for transactions and Easysearch for full-text search, syncing 1 million Chinese documents via Logstash; full sync completes in 5.5 minutes, incremental changes propagate within 30 seconds, and search latency drops 12x compared to SQL LIKE while maintaining identical recall.

CDCChinese tokenizationDameng DM8
0 likes · 20 min read
Dameng DM8 + Easysearch Full-Text Search Validation: 12x Speedup on 1M Chinese Docs
Mingyi World Elasticsearch
Mingyi World Elasticsearch
Sep 28, 2026 · Databases

Dameng DM8 + Easysearch: 17x Faster Full-Text Search via Logstash Sync

This article details integrating Dameng DM8 database with Easysearch 2.4.0 via Logstash 9.0 for full-text search, covering version compatibility fixes, JDBC driver configuration, logical deletion handling, timestamp boundary issues, IK analyzer alignment for Chinese text, and benchmark results showing 17x faster queries with consistent recall versus LIKE.

Chinese tokenizationDM8Dameng
0 likes · 16 min read
Dameng DM8 + Easysearch: 17x Faster Full-Text Search via Logstash Sync
360 Zhihui Cloud Developer
360 Zhihui Cloud Developer
Apr 9, 2026 · Databases

Master PostgreSQL Full-Text Search: From Basics to Advanced Chinese Tokenization

This article explains PostgreSQL's native full‑text search, its core concepts of tsvector and tsquery, demonstrates how to use built‑in functions and operators, compares built‑in, zhparser, and pg_search extensions for Chinese tokenization, and provides best‑practice tips for indexing, triggers, and performance optimization.

BM25Chinese tokenizationPostgreSQL
0 likes · 14 min read
Master PostgreSQL Full-Text Search: From Basics to Advanced Chinese Tokenization
Wukong Talks Architecture
Wukong Talks Architecture
Mar 31, 2021 · Backend Development

How to Install and Use the IK Chinese Analyzer Plugin in Elasticsearch

This article explains why Elasticsearch's built‑in tokenizers struggle with Chinese text, introduces the IK analyzer plugin, provides step‑by‑step Docker and file‑based installation methods, shows how to configure custom dictionaries via Nginx, and demonstrates smart and max‑word tokenization queries.

Chinese tokenizationCustom DictionaryDocker
0 likes · 12 min read
How to Install and Use the IK Chinese Analyzer Plugin in Elasticsearch
Wukong Talks Architecture
Wukong Talks Architecture
Oct 9, 2020 · Big Data

Elasticsearch Fundamentals: Architecture, Indexing, Queries, Docker Setup, and Chinese Tokenization

This tutorial introduces Elasticsearch's core concepts, installation via Docker, index and document operations, query DSL, aggregations, and Chinese tokenization using the IK analyzer with custom dictionaries, providing step‑by‑step code examples for building a searchable log analysis stack.

Chinese tokenizationDockerElasticsearch
0 likes · 28 min read
Elasticsearch Fundamentals: Architecture, Indexing, Queries, Docker Setup, and Chinese Tokenization