Where Write Time Goes in Multi-Region Active-Active Databases
This article analyzes write latency in multi-region active-active databases like AWS Aurora DSQL, showing that local execution is only a small fraction of total latency while cross-region commit coordination dominates; it also examines optimistic concurrency control conflicts, hotspot contention, isolation level limitations, and criteria for when multi-region active-active truly benefits applications.
In a typical order request, a few SQL statements execute quickly yet the user still waits for "order success." Can a multi-region database that allows local reads and writes eliminate this wait? AWS Aurora DSQL (released July 2026, updated August 2026) separates transaction execution from commit coordination, providing a concrete case to dissect the problem. Understanding the design value requires following a write all the way to final confirmation, not stopping when SQL execution finishes.
Local Execution Still Requires a Commit Confirmation
According to AWS architecture documentation, a DSQL multi-region cluster executes the transaction in the region that receives the request. Cross-region coordination occurs only at commit time for modifying transactions; read-only transactions avoid this coordination cost entirely. "Local" here means within the database region, not zero network distance from the user's device.
Consider an order that queries a product, checks an account, and updates inventory. If every statement had to cross a region boundary to reach a remote primary, each business step would add a round trip. Moving execution locally shortens those round trips, but local completion does not equal final commit confirmation. The application must still wait for the commit result before promising the user that the order is placed.
This explains why SQL monitoring may look great while the API latency does not improve proportionally. Connection wait, application computation, database execution, commit confirmation, and retries after failure all lie on the user's wait path. Looking only at statement latency misses the later stages.
A concrete hypothetical breakdown for a single serial request: other overhead 10 ms, local execution 10 ms, commit phase 80 ms, total 100 ms. If local execution is reduced to 2 ms (an 80% drop) while the other two stages stay unchanged, total latency becomes 92 ms — only an 8% improvement. These numbers are not actual DSQL measurements nor latency promises for any specific region pair.
Figure 1: Hypothetical single request with 10 ms other overhead, 80 ms commit unchanged, local execution reduced from 10 ms to 2 ms, total latency drops from 100 ms to 92 ms. Bars share zero baseline; not actual measurements, not P99 addition.
The example only adds non-overlapping stages within the same request. You cannot simply sum P99 values from separate monitoring panels to get the API P99, because the slowest requests in each stage may not be the same requests. Real evaluation requires request tracing to align timings per full request and then computing percentiles.
Therefore, "multi-region active-active" first changes which regions can accept reads and writes and which stages require coordination. It does not automatically answer whether a given transaction can meet a response target. For long-distance deployments, measuring the commit phase is especially important; for businesses with long call chains, reducing serial calls outside the database may be more effective.
More Entry Points, Why Not More Inventory to Sell
If different users modify different orders, multiple regions can share the write load. But when all users target the last unit of a hot product, the number of write entry points increases while the conflicting business object remains a single row.
DSQL uses optimistic concurrency control (OCC), checking conflicts at commit. The current documentation states that concurrent modifications of the same row may cause one transaction to fail, requiring the application to handle retries. Absence of lock waits does not mean conflicting modifications can all succeed.
The last-inventory scenario illustrates this: two transactions each see stock available, then race to update the same inventory row. One commits first; the other is rejected. On retry, the latter must re-read inventory and re-evaluate whether the order can proceed. Simply resubmitting the old computation discards the business meaning of the retry.
Figure 2: Two entry points facing the same inventory item. Scenario illustrates that shared objects don't increase with more entry points; does not imply database commits via physical contention.
Under high conflict, the queries and computations that precede failure still consume resources; retries generate new requests. Unbounded, immediate retries without backoff can intensify contention. Validation must simultaneously examine conflict rate, attempts per successful transaction, and final response time. Counting only incoming write requests can mistake repeated work for throughput gains.
Randomizing order primary keys disperses order writes but does not eliminate contention on the inventory row. Sharding inventory into multiple quota shards has costs: each shard can only commit within its allocated quota, cross-shard rebalancing requires extra coordination, and total stock may remain while a specific shard is temporarily sold out. Whether dispersing the hotspot is worthwhile depends on whether the business can accept such allocation constraints.
Selection Still Missing One Question: Which Results Must Never Coexist
Beyond same-row contention, cross-row rules must be checked. DSQL documentation describes its isolation level as equivalent to PostgreSQL's Repeatable Read. Snapshot isolation cannot be directly treated as serializable; replica consistency does not mean arbitrary business constraints automatically hold.
For example, suppose a shift requires at least one person on duty. Persons A and B each occupy a row. Two concurrent transactions both read that both are on duty, then each updates their own row to off-duty. If the implementation relies only on normal snapshot reads, each writes its own row — this write skew is a risk that needs explicit verification. Both updates appear reasonable individually, but together they may leave the shift empty.
One approach is to make exit operations for the same shift contend on a common constraint record, checking and updating related status within the same transaction. This maps the rule onto conflict detection, but the shift record becomes a new contention point. Another candidate is explicitly declaring that read rows participate in conflict checking; DSQL's SELECT FOR UPDATE uses optimistic checking at commit and cannot be treated like traditional "read and hold lock" semantics. Any solution must cover all modification paths, including new hires and batch operations.
With this analysis, the selection conclusion becomes clearer. Businesses with heavy intra-region reads, dispersed write objects, and applications that correctly handle transaction retries are better positioned to benefit from local execution. Businesses that concentrate on deducting the same balance or inventory must first solve hotspot and constraint expression; adding regions alone will not eliminate competition. Businesses serving only one region with a tight response budget have no reason to add coordination distance just for "two places can write."
Final validation must preserve real key distributions, not rely solely on random keys for load testing; it must test normal commits, conflicts, lost commit receipts, and region unavailability. When a receipt is lost, first verify the business outcome — do not treat an unknown result as uncommitted. Traffic switching, connections, and dependent services outside the database must also work end-to-end. Only when business rules and response targets are met under these conditions does active-active deliver its promised value.
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
IT Architects Alliance
Discussion and exchange on system, internet, large‑scale distributed, high‑availability, and high‑performance architectures, as well as big data, machine learning, AI, and architecture adjustments with internet technologies. Includes real‑world large‑scale architecture case studies. Open to architects who have ideas and enjoy sharing.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
