Databases 20 min read

From Backup to Real-Time: Automating Data Recovery at 10M QPS

This article traces the evolution of data recovery from manual backup restoration through log replay to automated replica switching, explaining RPO/RTO trade-offs, and details how 10M QPS systems implement shard-level automatic rebuild with throttling, verification, and continuous recovery as a runtime capability.

Random Bulletin
Random Bulletin
Random Bulletin
From Backup to Real-Time: Automating Data Recovery at 10M QPS

Introduction: Two Failure Stories

The article opens with two contrasting scenarios. In the first, an e-commerce DBA is woken at 2 AM: the primary order database crashes, the latest full backup fails checksum validation, and the team falls back to a two-day-old backup plus binlog replay — taking 1.5 hours and costing 3× peak revenue per minute. In the second, a similar primary failure triggers an automatic replica takeover within 30 seconds, with near-zero business impact. The difference is not luck but how far each organization has evolved its recovery architecture.

RPO and RTO: The Two Independent Axes

RPO (Recovery Point Objective) answers "how much data can we lose" measured in time. RTO (Recovery Time Objective) answers "how long is service unavailable" from failure to restored availability. They are independent: fast recovery ≠ low data loss, and low data loss ≠ fast recovery.

A common pitfall is conflating them. Daily full backup + log replay can achieve near-zero RPO (if logs are intact) but RTO remains hours because restore is slow. Primary-replica failover gives second-level RTO, but replication lag means the last few milliseconds of writes may be lost (non-zero RPO). Concrete numbers: daily backup only → RPO up to 24 hours, RTO 4–12 hours; backup + complete log archive → RPO minutes, RTO still hours; real-time replica → RTO seconds, RPO = replication lag.

Different business lines tolerate different targets. Internal reporting may accept daily RPO and hourly RTO; payment core demands zero RPO and second-level RTO. Architects must first define explicit RPO/RTO per data class, then choose mechanisms. Without targets, solutions are either over-engineered or useless.

At 10M QPS, one second of RTO affects 10 million requests, many already queued and unreplayable. Recovery ceases to be a post-failure cleanup and becomes a core variable that determines incident severity.

Backup Recovery: Classic but Slow

The standard pattern: periodic full backup (daily/weekly) plus incremental backups, then on failure restore latest full, apply incrementals, replay logs to failure point. The bottleneck is sequential I/O during restore. On commodity SSD, restore throughput is ~hundreds of GB/hour; a 10 TB instance can take a day or more. Every step is serial (full → incremental → log replay), so total time is incompressible.

Hidden risk: the backup set itself may be unusable — corrupted files, inconsistent snapshot at backup time, or backup stored in the same AZ as primary. Industry rule: a backup not verified by actual restore equals no backup . Backups must undergo regular restore drills and be stored across AZs/regions.

Why keep backups if replicas exist? Logical errors (e.g., accidental DROP TABLE) replicate faithfully to all replicas. Only a pre-error snapshot (backup) can recover. Thus backup + PITR remains essential for data correctness.

Log Replay (PITR): Restore to Any Second

Databases write every change to an append-only log (redo log, binlog, WAL). Logs are fast to write and contain full change information. PITR (Point-in-Time Recovery) uses an earlier full backup as base, then replays all logs from backup time to target timestamp, reconstructing state at that exact second. This shrinks RPO from "backup interval" to "log durability lag".

However, PITR does not solve RTO. Replaying a day's logs can take hours. Standard uses: (1) recover a dropped table or corrupted data to a point before the error for comparison/rescue; (2) on a standby that has already taken over traffic, replay logs slowly while primary is down, so RTO depends only on failover time.

Critical engineering detail: the log chain must be unbroken. Any missing segment collapses PITR from "any second" to "up to the gap". Log archiving therefore requires the same multi-copy, checksummed reliability as the data itself.

Replica Switching: Move Recovery to Runtime

To achieve second-level RTO, keep a near-real-time replica continuously applying primary's log stream. On primary failure, redirect traffic to replica — RTO = detection + cutover (seconds).

Two hard boundaries:

Replication lag = RPO. Replica always trails by some log volume. Zero RPO requires primary to wait for at least one replica acknowledgment (semi-sync or multi-primary consensus), adding a network round-trip per write — latency and throughput penalty. At 10M QPS, strong sync is enabled only for a few critical tables; others accept millisecond-level potential loss.

Replicas cannot prevent logical errors. A mistaken bulk update or DROP TABLE replicates instantly. Healthy architecture uses replicas for availability, backups+PITR for correctness.

At 10M QPS with thousands of shards, each shard must have replicas and automatic failover. Recovery shifts from a one-off emergency procedure to a continuous, distributed system capability — the subject of the next section.

10M QPS Rebuild: From Manual Ops to Distributed System Capability

Context: thousands of shards, each multi-replica. Single-shard failure (disk, host) is a daily occurrence. The system must automatically rebuild the missing replica without human intervention and without impacting live traffic.

Standard rebuild flow (four steps):

Clone baseline data from a healthy replica or storage snapshot.

Attach to log stream and catch up on incremental changes.

Verify data correctness (checksum, row count, sampling).

Register replica into serving group.

This mirrors PITR's "base + log" but is distributed, automatic, and continuous.

Three new problems at 10M QPS scale:

Rebuild storm. A single host runs dozens of shard replicas. Host failure triggers dozens of concurrent rebuilds pulling full data from the same source, saturating its network/IO and degrading live traffic. Mitigations: priority queue per shard, limit concurrent rebuilds per source node, isolate rebuild traffic (dedicated NIC, separate IO queues, or storage-layer snapshot clone instead of network copy).

Rebuild speed vs. second failure. If one shard takes 6 hours to rebuild and another replica fails in that window, data loss becomes real. Target: single-shard rebuild in minutes. Techniques: storage-layer snapshots (TB-scale clone in seconds), transfer only changed blocks, multi-threaded parallel replay.

Correctness during rebuild. The catching-up replica holds partial state and must not serve reads. Only after verification — per-shard checksum match, row-count match, sampled row-by-row comparison — can it join the serving group. Checksum consistency is a hard gate.

Together, these make recovery a scheduled, throttled, verified background distributed task . When failure strikes, no one runs emergency scripts.

Recovery Verification: Don't Trust "Restore Successful"

Logs may show success while data is silently wrong (e.g., missing secondary index, table stuck at unexpected timestamp). Verification has three layers:

Integrity checks (automated, low cost): row counts, per-shard checksums, boundary values (max ID, last update time). Catches most "half-restored" cases.

Business consistency checks (defined with product): sample key queries — e.g., can we see orders from last 10 minutes? Are state transitions valid? Part of regular drills.

Traffic validation : canary a small fraction of read-only traffic, monitor error rate and latency, then ramp. Requires read-write separation and shadow-traffic support.

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

replicationdata recoverybackuphigh QPSRPORTOpoint-in-time recoveryshard rebuild
Random Bulletin
Written by

Random Bulletin

17-year internet software developer specializing in AI applications, networking, architecture, and open source. Led the delivery of network services handling hundreds of millions of concurrent devices and tens of millions of QPS, and has three years of experience designing and building an agent platform. Follow to stay updated.

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.