Ordered Messages Aren't Enough: Preventing State Regressions with Versioning and State Machines

The article explains why message ordering alone fails to prevent business state regressions, detailing how version gaps and unordered consumption cause issues, and presents solutions using event versioning, state machines, and proper Kafka/RocketMQ configuration to ensure correct order processing per order ID while allowing parallelism across orders.

Niu Liu
Niu Liu
Niu Liu
Ordered Messages Aren't Enough: Preventing State Regressions with Versioning and State Machines

1. Message Not Lost but Order State Regressed

An order that has already shipped suddenly reverts to "awaiting payment." The root cause: messages for create, pay, and ship arrive in order, but the consumer throws them into a thread pool. A slow query on the create event delays its persistence, while pay and ship events complete first, overwriting the newer state with the older one. Messages arrive in order, but business effects do not. Simply enabling "ordered messages" without understanding the gaps will eventually cost debugging time.

Flowchart of ordered processing for same order and parallel execution for different order units
Flowchart of ordered processing for same order and parallel execution for different order units

Same order processed serially; different order units run in parallel. Version gaps only block the affected scope.

2. Only the Same Order Needs Ordering

Order A must pay before ship; Order B does not need to wait. Using order ID as the ordering key serializes events for the same order while allowing different orders to run in parallel across processing units. A single global queue would make all orders wait for the slowest one — not a good choice.

There are three distinct orderings: business generation order, broker write order, database commit order. A shared key only helps routing; it cannot decide business precedence across multiple producers. The author's approach: let the authoritative order writer generate an event version inside the same transaction that modifies the order, then send messages serially by that version. Avoid letting multiple services each add versions, and do not rely on machine timestamps to guess order.

Messaging System Comparison

RocketMQ 4.x : Same key routes to same queue; use ordered consumer listener. Guarantees queue-level order.

RocketMQ 5.x FIFO : FIFO topic, set message group; same group sends serially. Guarantees message-group order.

Kafka regular partition consumption : Stable key routes to same partition; enable producer idempotence; process serially. Guarantees partition log order.

RocketMQ 5.x ordered sending requires a single producer and serial sends. Kafka producer idempotence handles protocol retries but does not reorder business events from multiple producers. If the consumer then uses an unconstrained thread pool, neither side can save ordering.

3. Version Numbers Block Gaps, State Machines Block Chaos

Current order version is 7; only version 8 may proceed. If version 9 arrives, the missing 8 must be fetched first. Already-processed event IDs can be idempotently acknowledged; unfamiliar old versions must be verified, not blindly discarded as duplicates.

Below is abstract pseudocode assuming a complete continuous event stream with initial version 0. Subscriptions that filter out intermediate events cannot directly apply the +1 rule.

void onMessage(Event e) {
    try {
        db.transaction(() -> {
            Order o = orders.lockOrCreateInitial(e.orderId);
            if (inbox.exists(handler, e.id)) return;
            if (e.version != o.version + 1) throw new SequenceGap();
            stateMachine.requireAllowed(o.state, e.type);
            orders.applyAndAdvanceVersion(o, e);
            inbox.insertUnique(handler, e.id);
        });
        delivery.confirm(e); // confirm after business commit, or advance continuous completion offset
    } catch (SequenceGap gap) {
        delivery.pauseAndRetainUnprocessed();
        repair.fetchMissingEvents(e.orderId); // do not acknowledge, do not skip
    }
    // other exceptions go to framework retry; must not advance consumption progress
}

Business update, version advance, and deduplication record must be in the same transaction. Row locks or equivalent concurrency control prevent two consumers from reading the same old version. External actions need their own reliable delivery and idempotence; database transactions cannot cover remote calls.

If version 9 is stuck at the head while version 8 is behind, retrying 9 locally is useless. A repair process must fetch missing events from the authoritative event log and then resume consumption. Having only a pause button without a refill path causes a deadlock.

However, in practice the version-number approach is often not pragmatic; state machines are more commonly used to correct out-of-order transitions.

4. Two Operations That Easily Break Ordering

Moving failed messages to a retry queue while subsequent messages continue. If message 2 fails and message 3 takes effect, strict ordering is lost. Either block the affected order, or explicitly allow skipping with compensation — do not try to have both.

Kafka's pause/resume does not rewind the read position. Failed records and unprocessed records in the same batch must be retained, or the consumer must reposition to the earliest unprocessed offset; committed offsets must not skip over unfinished records.

Partition expansion changes the mapping for the same key. With modulo routing, changing the partition count can remap a key to a different partition while historical messages remain in the old partition. Both old and new partitions consume in parallel, breaking order. RocketMQ's queue selection by queue count has the same risk. Mitigation: pause related production, wait for old-route consumption to reach the switch boundary, then enable new routing; if production cannot stop, manage routing versions and business ownership migration. Adding machines is not the same as adding ordering units.

5. Don't Block All Business for the Sake of Ordering

Order lifecycles and account-internal events suit local ordering. Full snapshot cache refreshes can be simplified if the business allows higher versions to overwrite lower ones. Penny-accurate fund and inventory changes cannot arbitrarily drop intermediate steps.

Before launch, the author deliberately makes ship arrive before pay, creates duplicate deliveries, and then expands partitions under backlog. Acceptance criteria: no level-skipping effects, gaps trigger alerts, and recovery resumes after fix. Routine monitoring should track not only total backlog but also the oldest message wait time and blocked business keys.

Let the MQ deliver in order as much as possible; let the business refuse to act incorrectly when ordering fails.

Real production is far more complex: duplicate push idempotence, ghost event isolation, business compensation events, etc.

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

distributed systemsState MachineKafkaMessage QueueRocketMQOrdered MessagesEvent Versioning
Niu Liu
Written by

Niu Liu

A slightly rustic name 🤠 A tech veteran navigating the internet wave Hardcore tech: fixing all bugs and tough challenges

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.