Same-City Disaster Recovery: From Cold to Hot Backup at 10M QPS
This article details the evolution from cold to hot backup in same-city disaster recovery, covering fault domain analysis, RPO/RTO definitions, data replication strategies, capacity planning, traffic switching challenges, control plane resilience, and phased implementation approaches for systems handling ten million QPS.
At 2 a.m., the core switch in the primary data center begins restarting repeatedly. Monitoring first reports network jitter, then massive database connection timeouts; within minutes the business success rate drops from 99.99% to below 70%. The on‑call engineer opens the emergency playbook, which reassuringly states: “Switch to the same‑city disaster‑recovery data center if necessary.”
When actually executed, problems cascade: no one knows how far the standby database lags; a dependency was missed during last month’s capacity expansion; the traffic entry point can be redirected, but payment callbacks still point to the primary site; the ops platform itself runs in the failing data center. Worse, the primary site hasn’t fully lost power — intermittent connectivity makes both sites believe they should keep serving traffic.
Having a standby data center does not equal having same‑city disaster‑recovery capability. Facilities, servers, and dedicated lines are merely materials; the real determinant is whether data replication, capacity reserves, traffic scheduling, failure detection, and organizational coordination form a closed loop.
“From cold backup to hot backup” is not a simple equipment addition. Each step forward shortens recovery time but increases the number of simultaneously running states, raising data‑consistency risks, false‑failover risks, and daily operational costs. At ten‑million‑QPS scale these tensions amplify: the switch target grows from dozens of machines to thousands of instances; cache rebuilds can instantly overwhelm backends; a few minutes of backlog may correspond to billions of requests.
First, clarify what same‑city DR actually protects
“Same‑city” typically means two or more data centers within acceptable network latency, connected by metropolitan fiber or dedicated lines. It emphasizes low‑latency collaboration, not inherent independence. Two sites 30 km apart may share a substation, the same inbound fiber, or the same carrier aggregation node. They may occupy different buildings yet share unified authentication, configuration centers, and control planes. Judging isolation by straight‑line distance on a map easily overestimates true independence.
Map fault domains before drawing deployment diagrams
The first diagram for same‑city DR design should not be the application architecture but a fault‑domain map. You must answer layer by layer: when a component fails, which capabilities disappear simultaneously?
Common fault domains include at least:
Single‑machine and single‑disk failures — usually handled by local HA.
Rack, network partition, or power‑unit failures — require cross‑rack scheduling.
Entire data‑center, building, or campus failures — the primary focus of same‑city DR.
City‑level power, backbone network, or natural disasters — should not rely solely on same‑city solutions.
Human misoperation, bad releases, and data corruption — can propagate to all sites simultaneously.
The last category is especially easy to overlook. Hot backup can survive hardware and site failures, yet may replicate an erroneous deletion in real time. DR must work with backups, version rollback, and logical isolation — they cannot replace each other.
Bound the design with service‑level commitments
Any DR solution must answer two metrics:
RPO (Recovery Point Objective) — maximum tolerable data‑loss window.
RTO (Recovery Time Objective) — maximum allowable interruption from failure to recovery.
These cannot live only on a slide. Order creation, product search, recommendation features, and offline reports have vastly different tolerances. Writing “RPO = 0, RTO < 1 minute” for the whole system is usually both expensive and undeliverable.
Only when fault domains, RPO, and RTO are concrete can cold, warm, and hot backup be compared on the same scale.
Why cold backup is cheap — and why it often recovers slowly
Typical cold backup retains the facility, base resources, backup data, and deployment artifacts at the standby site. It carries no production traffic; some compute resources aren’t even started. On failure, capacity is requested, data restored, applications started, and entry points switched.
Advantages are clear: low standby resource utilization, low long‑term cost; no simultaneous writes, so fewer dual‑active conflicts. For small systems where recovery time measured in hours is acceptable, cold backup may be rational.
The problem: cold backup pushes massive work to the moment of failure.
Cold‑backup RTO is a chain of serial delays
A cold‑backup recovery typically involves these steps:
Confirm the primary site cannot recover within the target window.
Request or start compute, network, and storage resources.
Locate the latest usable backup and verify integrity.
Restore full data, then replay incremental logs.
Deploy applications, load configurations, register services.
Validate core paths, switch external traffic.
Process backlogged tasks, cache refill, downstream retries.
Total recovery time is not the maximum of these steps but the sum along the critical path. If backup restore takes 90 minutes, even instant app startup and traffic switching cannot meet a 15‑minute RTO.
A common drill pitfall: testing only “can the backup be restored” without testing “can the restored system serve traffic.” Successfully mounting data files does not mean the business works. Database accounts may have expired, certificates may be out of sync, external whitelists may list only the primary egress IP — small issues that break the RTO promise.
Data scale magnifies cold‑backup weaknesses
Assume 500 TB must be restored at an effective 5 GB/s throughput. Ideal full read‑write alone takes nearly 28 hours. Reality adds small files, checksums, throttling, storage contention, and retries. Adding more restore nodes helps only in certain stages.
Ten‑million‑QPS systems add further pain: vast amounts of caches, indexes, features, and derived state surround the primary data. Even if the primary database recovers, search indexes not caught up, empty caches, and inconsistent message offsets can crush the system with rebuild traffic the moment it comes online.
Cold backup isn’t unusable — it just demands honest acceptance of its boundaries. It suits systems where RTO is measured in hours, resource‑recovery paths are stable, and data volume is manageable. It also serves as a last‑resort recovery layer beyond hot backup.
Moving from cold to warm: advance the recovery actions
Warm backup has no single industry definition. In practice it means: the standby site runs key infrastructure and data replicas continuously; applications stay scaled‑down or on standby; on failure they scale up, validate, and cut traffic.
Compared to cold backup, warm backup simply moves failure‑time work to peacetime:
Continuous async replication avoids massive full restores during incidents.
Network, certificates, DNS, and security policies are pre‑provisioned.
Applications run at reduced spec, continuously passing health checks.
Release pipelines cover both primary and standby sites.
Core dependencies are probed regularly, not discovered at cutover.
“Having a replica” ≠ “replica can take over”
Async replication typically shrinks RPO from hours to seconds or minutes, but introduces replication lag and state‑judgment issues. Monitoring must go beyond “replication process alive” to track:
Gap between received log offset and primary commit offset.
Replica replay lag, not just network receive lag.
Silent stalls in the replication pipeline.
Whether the replica passes consistency checks.
Upstream data version compatibility with the application version.
Example: the standby database has received the latest logs, but the replay thread is blocked by a large transaction, leaving readable data 20 minutes behind. If monitoring only reports “replication connection healthy,” cutover will cause business state rollback.
Scaled‑down operation must answer the capacity‑ramp question
The warm standby may hold only 20%–40% of production capacity. On failure it must scale to full size. Design cannot just say “auto‑scale”; you must measure:
Time from scale‑out trigger to new instances ready.
Cache‑warm‑up duration.
Message‑queue catch‑up time.
Downstream dependency connection‑establishment latency.
If target RTO is 10 minutes but the P95 scale‑and‑warm time is 18 minutes, the plan fails on paper. Options: increase standing capacity, accept deeper degradation, or renegotiate RTO.
Warm backup’s greatest value: it lets teams build cross‑site data, release, and drill capabilities without taking on the full complexity of hot backup. Many systems should solidify warm backup first, then decide if hot backup is truly needed.
Hot backup isn’t “everything running” — it’s “provably ready to take over at any moment”
The hot standby usually maintains full or near‑full running capability: continuous data sync, application versions and configs identical to primary. It may carry zero user traffic, or a small shadow/read‑only/proportional slice of production traffic.
Hot backup pursues shorter RTO and smaller RPO, but “hot” is easily misunderstood. Servers powered on only proves processes are alive. True hot backup must satisfy at least four conditions:
Data state meets the business‑defined RPO.
Compute and dependency capacity can absorb traffic within the target time.
Traffic entry points have an operable, reversible switch path.
The team proves the whole process works via periodic drills.
Primary‑backup hot standby ≠ active‑active
In primary‑backup hot standby, the primary handles reads/writes; the standby stays hot and promotes on failure. The write path is relatively clear, but the cutover must handle role change and old‑primary isolation.
Active‑active (or multi‑active) serves production traffic from multiple sites simultaneously. Resource utilization is higher; standby issues surface during normal operation; single‑site traffic migration is smoother. The cost: significantly more complex data writes, session affinity, cross‑site calls, and conflict resolution.
The yardstick for hot backup is not whether resources are powered on, but whether they continuously meet takeover conditions — and whether that conclusion has real‑time evidence.
Let the standby shoulder real but controlled work
A standby with zero production traffic silently rots. A safer approach: let it handle a slice of real work — read‑only queries, internal‑user traffic, shadow requests, or discardable async tasks.
This continuously exposes certificate, routing, dependency, and capacity issues. Side effects must be contained: shadow requests must not double‑charge or send duplicate messages; read‑only traffic must not turn cross‑site reads into a new latency source. Whether to carry production traffic depends on business side‑effects and data model.
Data replication: the closer RPO is to zero, the harder the write path
Same‑city low latency makes synchronous replication feasible, but “feasible” ≠ “free.” Assume 1 ms one‑way latency; a synchronous write adds at least one cross‑site round‑trip, remote storage commit, and protocol overhead. Per‑transaction impact seems small, but multiplied across multiple serial write layers, tail latency rises sharply.
Ten‑million‑QPS systems cannot simply mandate synchronous replication for all data. You must stratify by data semantics.
Three common replication strategies
Synchronous / quorum commit. Write succeeds only after cross‑fault‑domain replica acknowledges. RPO approaches zero. Suits ledger accounts, order master state — data that cannot be easily compensated. Cost: higher write latency; cross‑site network issues can reduce availability.
Asynchronous replication. Primary commits locally and responds immediately; logs ship to standby afterward. Keeps write latency low and primary availability high, but failure may lose un‑replicated data. Fits replayable, compensable, or slightly rollback‑tolerant data.
Business‑level rebuild. Derived data (search indexes, recommendation features, caches) need not be replicated verbatim; they can be regenerated from primary data, message logs, or object storage. Saves replication cost, but rebuild time must be counted into RTO.
Synchronous replication can still suffer common‑mode failures
If two data replicas span sites but the arbiter and primary replica sit in the same site, a primary‑site network partition may prevent the standby from forming a quorum. Conversely, allowing either side to independently write for availability risks dual‑master and data divergence.
Majority protocols solve node consensus, not fault‑domain planning. Replica count, voting weights, and deployment locations must be co‑designed. Symmetric two‑site deployments often hit the even‑vote problem; a third arbitration point helps, but its own availability, latency, and security must be factored in.
Fence the old primary before promoting the new
During a primary‑site network partition, the most dangerous move is promoting the standby while the old primary still accepts writes. The correct sequence is typically:
Confirm the old primary cannot write — via storage fencing, network isolation, leases, or power control.
Verify standby replica offset and data integrity.
Promote the new primary and freeze topology changes.
Switch application connections and ingress traffic.
When the old primary rejoins, run data validation or full rebuild.
Such isolation actions are called fencing . They are hard constraints against split‑brain, not something to leave to “it should be down” human judgment.
Traffic switching: the real difficulty is moving every entry point together
Many playbooks reduce switching to changing one DNS record. Real systems simultaneously have public DNS, client long‑connections, dedicated‑line entrances, API gateways, message callbacks, scheduled jobs, and internal service discovery. Switching only a subset creates a “half‑moved” state.
List every entry point first
You must catalog each traffic type’s control point, convergence time, and rollback method:
Public user traffic — controlled by DNS, Anycast, or global load balancer.
Mobile/device long‑connections — subject to client reconnection logic.
Enterprise dedicated lines — may require customer‑side routing changes.
Third‑party callbacks — depend on external platform whitelists and address config.
Internal RPC — relies on service discovery, gateways, cross‑site routing.
Message consumers, batch jobs, schedulers — may lack a network entrance but keep generating writes.
Setting DNS TTL to 30 seconds does not guarantee all traffic shifts in 30 seconds. Recursive resolvers may cache longer; clients may reuse connections for extended periods; some networks ignore excessively low TTL. Switch design must allow old and new entry points to coexist for a period, while ensuring the old entrance stops writing to the wrong data primary.
Phased rollout is more controllable than big‑bang
The hot standby should already handle small traffic in peacetime. During failure cutover, ramp up by region, tenant, user hash, or business priority:
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
Random Bulletin
17-year internet software developer specializing in AI applications, networking, architecture, and open source. Led the delivery of network services handling hundreds of millions of concurrent devices and tens of millions of QPS, and has three years of experience designing and building an agent platform. Follow to stay updated.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
