Exactly Once: What It Guarantees, Why Payments Double-Charge, and How to Fix It
This article explains that Exactly Once semantics in message queues like Kafka only guarantee exactly-once processing within the messaging system's transaction boundary, not for external side effects like payment APIs, and provides practical patterns using database unique constraints, idempotency keys, and timeout handling to achieve true end-to-end exactly-once behavior.
1. Message Redelivered, Money Deducted Twice
If someone claims "enabled Exactly Once, business will never duplicate," I would walk through this scenario: consumer calls payment API successfully, crashes before committing consumption offset. On restart the message arrives again, code charges again.
The problem: the payment API did not participate in the MQ transaction, so it cannot roll back with it.
Exactly Once, in my understanding, means: a specific result, within a specific scope, takes effect only once. It does not mean every line of code executes only once.
Kafka internal transaction and external payment are two separate scopes that must be handled independently.
2. Kafka Is Powerful, But Has Boundaries
Kafka idempotent producer solves duplicate appends caused by protocol retries, not deduplication by message content. If the application actively calls send() twice, even with identical content, it may produce two valid records. (KafkaProducer documentation)
Kafka transactions further commit output records together with input consumption offsets, suitable for "read from one Topic, process, write to another Topic." Downstream must use read_committed isolation level to avoid reading aborted transaction outputs.
However, putting database updates or payment.charge() inside the processing function does not automatically enlist them in the Kafka transaction.
Recovery detail: aborting a transaction does not automatically roll back the consumer's already-advanced read position. The failed batch must be reprocessed, or the consumer thread must reposition; commit timeout follows API contract for recovery, cannot simply skip to next batch. (Kafka consumer position)
Kafka Idempotent Producer : Primarily handles protocol retry duplicate appends; does not cover application-level resends, duplicate business.
Kafka Transactions : Primarily handles Kafka output and consumption offset atomic commit; does not cover external databases, payment APIs.
RocketMQ Transaction Messages : Primarily handles upstream local transaction and message send eventual consistency; does not cover downstream consumption idempotency.
RabbitMQ Acknowledgments : Primarily handles publish and consume acknowledgments, support recovery; does not cover unified transaction across database and MQ.
RocketMQ transaction messages and RabbitMQ publish confirms cannot be directly interpreted as "business executes only once." (RocketMQ transaction scope, RabbitMQ acknowledgment mechanism)
3. Fix the Database: Deduplication Record and Business in Same Transaction
For ordinary business consumption that updates a database, I trust stable event IDs, unique constraints, and local transactions more.
Event redelivery, failure resend, and replay all retain the original eventId. The deduplication key can be "stable business processor identifier + eventId"; do not change processor name on each deploy, and do not rely on broker message IDs that may change on resend.
// Pseudo-code: unique conflict returns false, other DB errors re-thrown.
void onMessage(Event e) {
db.transaction(() -> {
if (!inbox.insertIfAbsent(handlerId, e.id)) return;
ledger.apply(e); // same database, same transaction as deduplication record
});
mq.confirm(e); // confirm only after commit success; failure left to framework recovery
}Database committed, acknowledgment lost — on redelivery the unique key blocks second modification. Exit before commit, transaction rolls back, next attempt executes normally. Do not fake atomic deduplication with "check-then-insert," and do not write Redis marker first then update database.
Business operations also need their own unique identifiers. Same payment intent generating two different event IDs cannot be blocked by event-level deduplication alone; pre-authorization, supplemental charge, and refund cannot all use order number as the same operation. Same key but different critical parameters (e.g., amount) should raise conflict.
Idempotency does not replace concurrency control. Different events modifying balance concurrently still require atomic updates, row locks, or version conditions. Kafka commits partition progress; under concurrent processing only continuous completed positions can be advanced.
4. Call Payment API: Clarify Result After Timeout
I first create a persistent payment task in a local transaction, with a unique constraint on business operation ID, commit then acknowledge message. Background task calls the channel, always using the same idempotency key.
Thus duplicate messages find the same task instead of creating a new charge. Background scanner recovers unfinished tasks; state must distinguish "taken" from "payment succeeded."
Request timeout enters pending-confirmation state; query result per channel protocol or retry with original idempotency key. Duplicate callbacks must not revert success state.
Idempotency keys may have retention limits. For example, Stripe's keys are not retained permanently; if result unknown after window expires, stop blind retries, switch to query, reconciliation, or manual handling — never change key and hope. (Payment idempotency conventions)
Timeout means unknown result, not that the other side didn't charge. If the channel lacks reliable idempotency and cannot verify results, local deduplication table alone cannot promise end-to-end single charge.
5. Acceptance Tests Look at Business Results
I test three things: kill process immediately after database commit; two consumers processing same event simultaneously; payment completed but response lost. Then verify ledger, balance, and channel transactions — not log prints of "success."
Deduplication records must cover the agreed retry, resend, and replay windows; long-term fund operations also need business ledger constraints. Monitor long-pending tasks and reconciliation discrepancies.
Use Kafka transactions for Kafka-to-Kafka pipelines; for database writes, implement idempotency and local transactions properly; for cross-payment systems, complete idempotency protocol and result verification. Messages may come again, but the same business must not take effect twice without reason.
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
Niu Liu
A slightly rustic name 🤠 A tech veteran navigating the internet wave Hardcore tech: fixing all bugs and tough challenges
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
