Real-Time Data Protection at 10M QPS: From Post-Incident Recovery to Instant Damage Control
This article details how high-throughput systems shift data protection from reactive backups to real-time guardrails, covering write-path constraints, CDC-based anomaly detection, reversible change state machines, object-level rollback with downstream compensation, and a phased evolution path to automate damage control at ten million QPS.
Why Backup Is Always a Step Behind
Traditional data protection relies on scheduled full backups, continuous log shipping, and post-failure recovery playbooks. This works for disk failures or instance outages where the fault boundary is clear. Logical errors are different: a malformed batch statement commits successfully, a buggy service version writes semantically wrong but syntactically valid data, or stolen credentials execute deletes that pass transaction checks. Storage systems faithfully replicate these errors, and replicas even accelerate their spread.
A timeline illustrates the passive nature of after-the-fact recovery. The most expensive gaps are three time deltas: detection latency, decision latency, and recovery latency. At millions of writes per minute, even a tiny fraction of corrupting requests produces massive dirty records in minutes. Worse, the bad data has already been consumed — message queues triggered refunds, search indexes updated statuses, risk systems adjusted limits, recommendation engines recomputed features. Restoring the primary table solves only part of the problem.
Therefore, data protection must answer five questions:
How quickly can an erroneous change be identified?
Can further writes be automatically throttled once detected?
Can committed data be rolled back at the business-object level instead of a full database rewind?
Can downstream side effects be pinpointed and compensated?
Who proves the recovery result is correct?
These questions push data protection from "keep a copy" to "control every high-risk change."
Define Protection Objects Before Talking Real-Time
"Protect all data in real time" sounds safe but becomes an expensive, vague goal. Data varies in value, update frequency, rebuildability, and regulatory requirements. The article classifies data into three tiers:
Authoritative state — decides business facts; needs strictest real-time protection.
Derived state — can be rebuilt from trusted sources; only needs to know which partitions require rebuild.
Temporary data — low recovery value; giving it the same protection level as financial data adds cost and ops complexity.
Two often-confused metrics are defined:
RPO (Recovery Point Objective) — maximum acceptable data loss window (e.g., 5 minutes).
RTO (Recovery Time Objective) — time from incident to service restoration (e.g., 30 minutes); does not guarantee full consistency.
Real-time protection does not mean RPO=RTO=0. Strict zero often requires cross-fault-domain synchronous acknowledgment, raising latency, cost, and failure coupling. A practical approach sets protection targets per data tier.
Protection levels must bind to business loss models, not be decided by storage teams alone. Losing 10,000 ad-impression logs affects only stats precision; losing 10,000 account balances becomes a financial incident. Only after quantifying loss can you justify detection speed, interception strength, and recovery granularity.
Real-Time Protection Must Sit Beside the Write Path
Backup lives after the write. Real-time protection must be close to the write path but cannot stuff all complex judgments into the database transaction, or the protection system itself becomes a latency and availability bottleneck.
A common architecture splits into four layers:
Synchronous path — only deterministic, low-cost checks: per-delete ceiling, tenant write-rate limits, object version, operator identity, idempotency keys. Must complete in milliseconds with explicit failure policies.
Asynchronous detection (sidecar) — CDC turns committed changes into events; detection system builds baselines per table, tenant, operation type, and business dimension. Anomalies in delete rate, field distribution, or object coverage trigger freeze, read-only, degradation, or secondary confirmation.
Control plane — executes responsive actions: freeze writes for a tenant/logical table/shard; block high-risk ops (e.g., unconditional bulk delete); divert writes to append-only staging log; reduce batch concurrency or batch size; revoke write permission for a specific release version; tag subsequent changes as isolated to stop propagation.
Recovery plane — orchestrates rollback and verification.
A common pitfall: sidecar detection only finds already-committed errors. If response still relies on humans, you merely shrink "hours later" to "minutes later" — not true real-time damage control. Detection output must connect to an executable control plane.
Finer action scope reduces collateral damage. Global read-only works but costs the most; freezing a single tenant's logical table avoids pausing all users.
Turn Dangerous Operations into Controllable State Machines
Many data incidents stem from "legal but dangerous" ops: historical cleanup, schema migrations, bulk migrations, state backfills. These happen routinely and cannot be simply banned. A more reliable approach converts them from one-shot commands into state machines with budgets, observation windows, and pause capability.
Four key mechanisms underpin this state machine:
Change Budget
Each operation declares maximum affected objects, tenants, shards, and duration. The executor draws quota batch by batch. Actual impact exceeding declared values triggers automatic pause. Budget dimensions must include object value, tenant concentration, field sensitivity, and downstream count — not just row count. Deleting 100k expired temp records differs entirely from modifying 100k active account settlement statuses.
Preview & Diff
High-risk ops first run a read-only pre-check, emitting samples, counts, distributions, and estimated side effects. The pre-check result is versioned; at execution time the system verifies the data baseline hasn't shifted, preventing "preview hit 100 rows, execution expanded to 1 million rows."
Small-Batch Execution
One giant uninterruptible transaction becomes many auditable micro-batches. Each batch carries a unique ID, input range, executor, start time, result summary, and reverse metadata. Gaps between batches give the detection system time to evaluate business metrics.
Auto-Pause
The system defines explicit signals that trigger pause: delete rate exceeding historical percentile, rising failure rate, widening reconciliation gaps, downstream consumption latency spikes. Pause is not rollback; it merely stops further damage buildup, buying time for judgment.
Practical guardrails let dangerous ops stop, inspect, and retreat at any moment.
Reversibility Is Not "Write an Opposite SQL"
Many teams equate rollback with executing a reverse statement. For simple field updates this may work; for concurrent writes, cross-table relationships, and external side effects, a reverse statement easily overwrites legitimate changes made after the incident.
Example: a batch job changes account status from A to B. During the incident, some users legitimately move their status to C via normal flows. Blindly reverting all B back to A destroys the valid C states.
Object-level recovery requires preserving rich change context:
Object primary key and tenant boundary
Before and after versions
Change source, operator, and release version
Business request ID, batch ID, idempotency key
Pre-change critical field values
Related events and downstream delivery offsets
Database commit time and logical sequence
With this context, rollback uses conditional updates: restore old value only if current version still equals the erroneous version; if the object later underwent legitimate changes, it enters a conflict queue for business-rule merge or manual handling.
Common reversible mechanisms each have applicability boundaries:
Real systems combine them: database logs for broad safety net, critical tables keep row-level before-images, deletes enter a recycled state first, cross-system actions record replayable events and compensation steps.
Recovery cannot skip downstream. After primary restore, at least four derived-state categories need handling:
Caches holding stale values — invalidate by version.
Search indexes that consumed bad events — targeted rebuild.
Message consumers that executed external actions — compensate or manual review.
Data warehouses and feature platforms that produced wrong results — mark contamination window and recompute.
Without a unified change ID threading through write, CDC, messaging, and derived tasks, answering "which downstream systems were hit by this bad write?" becomes nearly impossible. The protection system must propagate change identity end-to-end.
Detect Anomalies, Don't Just Watch Error Codes
The scariest trait of logical data incidents: they often emit zero error codes. SQL succeeds, transaction commits, replication is healthy, APIs return 200 — technical metrics look perfectly calm.
Detection must shift from system health to data behavior. Start by building multi-dimensional baselines for each change class:
A single static threshold cannot cover all tenants. A large tenant's normal batch may be hundreds of times a small tenant's peak. A robust approach combines three judgments:
Absolute ceiling — hard boundary that must never be crossed.
Relative baseline — compare current value against own historical distribution.
Business invariants — verify results obey conservation and state constraints.
Business invariants are the most valuable yet hardest to maintain. Examples: financial systems check debit-credit balance; inventory systems verify sellable + locked + sold consistency; order systems validate legal state transitions. These rules derive from the business model; generic monitoring platforms cannot infer them.
The detection system must also handle CDC event delay and out-of-order arrival. Events may arrive across partitions; a brief inconsistency should not trigger global freeze. Use a short aggregation window, merge by event time and business key, then assess whether anomaly persists. Window too long slows damage control; too short raises false positives — tune per object value.
A risk-grading example:
Risk grades do not aim to precisely predict accident probability. Their role is to bind actions to evidence, preventing all alerts from dumping into a single unhandled notification channel.
At 10M QPS the Protection System Must Avoid Its Own Avalanche
Traffic amplification brings new engineering constraints for real-time protection. The immediate issue is event volume. If every change copies full before-images and synchronously calls multiple detection services, write amplification, network overhead, and latency explode.
Hence the protection pipeline must separate synchronous and asynchronous work.
Synchronous Path — Minimal Deterministic Checks
Identity and permission validity
Object version match
Hard budget exceedance
Freeze/read-only policy hit
Audit tags completeness
These checks should use locally cached, versioned policy snapshots. Even if the control plane is briefly unavailable, the data plane continues operating on the last known good policy. Highest-risk ops can fail-closed; ordinary writes fail-open or degrade per business choice.
Asynchronous Path — Heavy Computation
Compute change rates and field distributions
Correlate multiple sources for business reconciliation
Store compressed before-images
Trace downstream propagation scope
Generate recovery plans and impact reports
CDC platform cannot be a single global queue. Shard by business domain and protection tier: P0 data gets dedicated resources and latency budget; derived data tolerates larger backlogs. Consumers must be idempotent; checkpoints must align with output commits to avoid duplicate compensations on protection-system restart.
Peak hours must never let the protection system back-pressure the main business. Pre-define degradation order: first drop low-value field sampling, then delay P2 verification, only then consider weakening P1. P0 audit and hard guardrails require isolated resources, not shared with analytical workloads.
Storage cost also tiers: recent hours keep fine-grained before-images for fast query; older data compresses to object storage; even older retains only audit summaries and database logs. Deleting expired protection data likewise needs audit and delayed deletion — don't let the protection store become a new single point of failure.
The protection pipeline must be more restrained than the protected pipeline: synchronous checks few and deterministic, asynchronous analysis comprehensive but degradable.
From Alert to Damage Control Requires a Closed Loop
If the system only surfaces anomalies, incident handling still relies on on-call engineers manually jumping across platforms. A complete loop connects evidence, action, recovery, and verification.
"Freeze impact scope" is easily overlooked. Upon detection, the system must immediately snapshot:
Logical start and end time of the anomaly
Involved business keys, shards, tenants, and fields
Related release versions, service accounts, and batch IDs
Message topics and downstream tasks already reached
Current checkpoints and available recovery points
Triggered automatic actions and their policy versions
Without this snapshot, data keeps changing during investigation; the team gets different impact numbers every few minutes, preventing consistent judgment.
Recovery strategy selects by impact scope:
Recovery completion does not mean immediate unfreeze. First run technical verification: row counts, checksums, version continuity, replication status. Then run business verification: balances, inventory, order states, cross-table relationships. Only after passing, gradually reopen writes by tenant or shard, observe, then expand.
This process should be codified as machine-executable recovery plans containing input scope, recovery actions, preconditions, verification items, and rollback fallbacks. On-call engineers approve and trigger; they don't improvise commands at the incident scene.
Drill Focus Is Not "Can Backup Decompress"
Many recovery drills stop at infrastructure: spin up instance, download backup, replay logs, confirm DB starts. That only proves files are readable, not that business can recover.
An effective data protection drill must answer seven questions:
Can a class of logical error be detected within the target time window?
Can automatic guardrails contain impact within the expected boundary?
Can affected objects be exhaustively enumerated?
Does object recovery avoid overwriting legitimate post-incident writes?
Can downstream cache, index, and message side effects be handled?
Can business invariants prove recovery correctness?
Will unfreeze cause re-consumption of the same bad event batch?
Start with small-scope fault injection: simulate bulk mis-delete, stale-version overwrite, duplicate events, field corruption in an isolated tenant. Injection must carry explicit tags; protection system should auto-detect, freeze, and generate impact report. Then execute the recovery plan and compare against target outcome.
Drill metrics must shift from "did it finish" to quantifiable data:
MTTM (Mean Time To Mitigate) deserves special attention. Traditional postmortems track MTTD and MTTR but lack a measure for "when did damage actually stop spreading." For high-velocity write systems, stopping error propagation first — even if recovery is slower — is usually more rational than chasing fast recovery while bad writes continue.
Drills also expose organizational gaps: who approves which ops? Does on-call have freeze authority? Do business teams know their invariants? Can security, DB, and app teams share a single impact scope? Unresolved in peacetime, these questions get paid for with time during incidents.
A Practical Evolution Roadmap
Real-time data protection is not a one-shot platform project. Chasing full before-images, full business invariants, and auto-rollback from day one yields an overweight system that teams bypass due to false positives and maintenance burden.
Advance in four phases:
Phase 1: Make Recovery Trustworthy
Inventory authoritative vs. derived data; define RPO/RTO for core tables. Verify backups, logs, and object versions actually restore, and run drills to the business verification layer. First solve "does the last line of defense really exist?"
Phase 2: Make High-Risk Operations Visible
Standardize recording of operator, release version, batch ID, business request ID. Bring bulk delete, migration, backfill, permission changes into audit. Establish CDC channels; start with impact analysis and alerting — not yet auto-intercept.
Phase 3: Add Fine-Grained Guardrails
Introduce budget, pre-check, small-batch execution, and auto-pause for high-risk ops. Narrow freeze scope to tenant, shard, operation type. Add business invariants for P0 data; enable soft delete or delayed cleanup for critical deletes.
Phase 4: Form Auto-Recovery Closed Loop
Codify object-level rollback and downstream compensation plans. On detection hit, auto-generate impact scope and suggested actions. Low-risk, high-evidence scenarios auto-rollback; high-risk retain human approval but system executes and verifies.
At each step, observe three side effects: write latency increase, false-positive interception rate, protection data storage growth. If guardrails frequently hurt legitimate traffic, teams will disable or circumvent them — no design survives that.
Measure data protection maturity by: how fast error spread stops, whether recovery scope is precise, and whether the result can be proven correct.
Moving from post-incident recovery to real-time protection essentially turns incident-handling experience into rules, state machines, and automated loops beside the write path. Small systems first solidify backup and restore; as traffic and business value grow, incrementally add audit, CDC, guardrails, and object-level recovery. Ten million QPS doesn't change the definition of data correctness, but it compresses the human reaction window and makes every error propagate faster to more replicas and downstreams.
Take this question back to your team: if a syntactically valid but semantically wrong batch write hits production right now, how many seconds until you detect it, and to what radius can you confine the damage?
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
Random Bulletin
17-year internet software developer specializing in AI applications, networking, architecture, and open source. Led the delivery of network services handling hundreds of millions of concurrent devices and tens of millions of QPS, and has three years of experience designing and building an agent platform. Follow to stay updated.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
