DiDi Interview Deep-Dive: Solving Elasticsearch-MySQL Data Consistency

This article analyzes four patterns for keeping Elasticsearch synchronized with MySQL — synchronous dual-write, message-queue async, CDC via Canal/binlog, and scheduled reconciliation — explaining why CDC with version-based deduplication and periodic checksums is the production-grade choice for eventual consistency at scale.

Java Architect Handbook
Java Architect Handbook
Java Architect Handbook
DiDi Interview Deep-Dive: Solving Elasticsearch-MySQL Data Consistency

Interview Focus Areas

Problem Framing : Recognize this as a distributed dual-write consistency problem — two storage systems cannot share a local transaction.

Solution Evolution & Trade-offs : Trace the progression from synchronous dual-write to MQ async to binlog listening; articulate what each solves and what new issues it introduces.

Production Depth : Demonstrate hands-on experience with message loss, out-of-order delivery, reconciliation, and full rebuilds.

Core Answer

Elasticsearch and the database target eventual consistency , not strong consistency. ES is a near-real-time (NRT) engine; writes become visible only after a ~1 second refresh, making strong consistency impossible from the start.

Four mainstream approaches, usually combined:

Synchronous Dual-Write : Consistency — Poor, prone to inconsistency; Performance — Poor; Business Intrusion — High; Complexity — Low; Verdict — ❌ Rarely used

MQ Async Sync : Consistency — Eventual; Performance — Good; Business Intrusion — Medium; Complexity — Medium; Verdict — ✅ Common

Binlog Listening (Canal/Debezium) : Consistency — Eventual; Performance — Good; Business Intrusion — None; Complexity — High; Verdict — ✅ Production Mainstream

Scheduled Reconciliation : Consistency — Safety Net; Performance — -; Business Intrusion — Low; Complexity — Low; Verdict — ✅ Mandatory Fallback

One-liner strategy: Binlog subscription as primary pipeline, version numbers to prevent reordering, scheduled reconciliation as safety net, accept second-level eventual consistency.

Deep Analysis

1. The Dual-Write Dilemma

Product data lives in MySQL; search runs on ES. Every product change must update both:

MySQL and ES dual-write flow
MySQL and ES dual-write flow

Neither order works atomically:

MySQL succeeds, ES fails → DB has new data, search returns stale data.

ES succeeds, MySQL rolls back → Search finds record, detail page returns 404.

This is the classic distributed dual-write problem. Strong distributed transactions (e.g., Seata XA) are overkill for search scenarios.

Baseline : Accept eventual consistency, shrink the inconsistency window to seconds, guarantee eventual convergence.

2. Synchronous Dual-Write: First Thought, First Discarded

Write MySQL then ES directly in business code. Problems:

High Coupling : Search logic leaks into business code; index changes require updates everywhere.

Poor Performance : API latency tied to ES write latency; ES hiccups stall order placement.

Still Inconsistent : Second-step failure has no automatic recovery.

Only merit: simple implementation. Acceptable for tiny, non-critical systems; not for production.

3. MQ Async: Decoupled, But New Pitfalls

MQ async sync to ES
MQ async sync to ES

Flow:

Producer : Business updates MySQL in local transaction, commits, then publishes change event to MQ.

Consumer : Dedicated sync service consumes events and writes ES.

Benefits : Business/search fully decoupled; API performance insulated from ES; MQ provides retry.

Two new pits interviewers probe:

Pit 1: Message Loss . MySQL commits, app crashes before MQ send → change lost forever. Solutions:

Local Message Table : Insert record into sync_message table in same transaction; scheduled job scans unsent rows and republishes.

Transactional Messages : RocketMQ's half-message + check-back mechanism guarantees atomicity between local transaction and message send.

Pit 2: Message Reordering . Price changed to 99 then 199; concurrent consumption writes 199 first, then 99 overwrites with stale price. Solutions:

Ordered Messages : Route by product ID so same product lands in same queue, preserving order.

Version Control : More fundamental; detailed next.

4. Canal Binlog Listening: Production Mainstream

MQ approach still embeds message publishing in business code — new tables forget to publish, multiple data sources bypass unified entry. Better: subscribe MySQL binlog directly .

Canal binlog sync to ES with reconciliation
Canal binlog sync to ES with reconciliation

Key points:

Canal (Alibaba open-source) masquerades as a MySQL replica (slave), uses dump protocol to receive ROW-format binlog increments. MySQL thinks it's replicating to a standby.

Zero Business Intrusion : Business only writes MySQL; sync logic entirely external. Schema changes need no business code changes.

No Data Loss : Binlog written only on transaction commit; Canal manages positions for exactly-once resume.

Buffer with MQ : Insert MQ between Canal and sync service for buffering and peak-shaving; pipeline becomes very stable.

Trade-offs: Longer ops chain (Canal cluster, MQ, sync service), latency shifts from milliseconds to seconds — acceptable for search.

International equivalent: Debezium , typically paired with Kafka; same CDC principle.

Canal via MQ to ES
Canal via MQ to ES

5. Version Numbers Prevent Reordering: The Elegant Trick

Regardless of pipeline, reordering is unavoidable. Best defense: ES External Version Control .

Attach monotonically increasing version (MySQL update timestamp or dedicated version column auto-incremented in transaction). On ES write:

IndexRequest request = new IndexRequest("product").id(event.getProductId())// External version control: ES only accepts writes with version >= current.version(event.getVersion()).versionType(VersionType.EXTERNAL_GTE).source(event.getDataJson(), XContentType.JSON);try {  client.index(request, RequestOptions.DEFAULT);} catch (EsRejectedExecutionException e) {  // Version conflict = stale message arrived late; skip — expected normal case  log.warn("Stale version message skipped: {}", event);}

Brilliance: Even if 199 arrives before 99, ES compares versions, sees 99 is older, rejects write. Stale data never overwrites fresh data. Reordering shifts from "prevention" to "immunity" — same idea as CAS/optimistic locking.

ES version control blocks stale messages
ES version control blocks stale messages

6. Safety Net: Reconciliation & Rebuild

All above only reduce inconsistency probability, never to zero. Production must have fallbacks:

Scheduled Reconciliation : Off-peak job samples rows, compares MySQL vs ES (update time or field checksum), repairs ES using DB as source of truth.

Full Rebuild : Major index restructuring or massive corruption — full reload from MySQL into new index, then alias swap. Ultimate "already broken" remedy.

Reconcile ES against MySQL
Reconcile ES against MySQL

High-Frequency Follow-Ups

MySQL succeeds, MQ send fails? Local message table (transactional insert + scheduled retry) or RocketMQ transactional messages (half-message + check-back). Core: bring message send into transaction guarantee.

Reordering causes stale overwrite? Two layers: ordered messages by business ID (reduce occurrence) + ES external version control (immune when it happens).

Why not strong consistency (distributed transaction)? ES NRT mechanism makes strong consistency impossible (1s refresh delay); search tolerates seconds; paying strong-consistency performance cost is unjustified. Consistency level must match business value.

Production already has dirty data — how to fix? Small scope: reconciliation script repairs per DB; large scope: full index rebuild + alias switch.

Common Interview Variants

"How did you handle dual-write inconsistency?" (Same question, different wording)

"Explain Canal's working principle." (Tests fake-slave binlog pull mechanism)

"How to keep cache and DB consistent?" (Same pattern: update DB then invalidate cache + delayed double-delete + fallback — answer by analogy)

"Describe ES write flow; why near-real-time?" (Pivot to refresh mechanism)

Memory Mnemonic

Dual-write unreliable, async into MQ; binlog most stable, version prevents reorder; reconciliation safety net, eventual consistency is fine.

Summary

The core logic chain: Define problem (distributed dual-write → only eventual consistency) → Show evolution (sync dual-write → MQ async → Canal binlog) → Add details (local message table prevents loss, version numbers prevent reorder, reconciliation catches leftovers). Walk this line fluently, then volunteer "Our production uses Canal + MQ + version control + reconciliation combo" — strong closing signal.

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

reconciliationElasticsearchdata consistencyMySQLmessage queueCanalinterview preparationversion controlCDCeventual consistency
Java Architect Handbook
Written by

Java Architect Handbook

Focused on Java interview questions and practical article sharing, covering algorithms, databases, Spring Boot, microservices, high concurrency, JVM, Docker containers, and ELK-related knowledge. Looking forward to progressing together with you.

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.