Cache Rebuild Evolution at 10M QPS: From Cold Start to Hot Standby

This article details the evolution of cache reconstruction strategies for ten-million-QPS systems, covering cold-start dangers, persistence recovery limits, tiered hot-data backfill, dual-cluster hot standby for second-level failover, shard-level background rebuild with singleflight, and multi-layer database protection during rebuilds, advocating making pre-warming a routine capability.

Random Bulletin
Random Bulletin
Random Bulletin
Cache Rebuild Evolution at 10M QPS: From Cold Start to Hot Standby

The destructive power of a cache failure lies not in data loss but in the backsource storm when hit rate drops to zero. A real incident at an e-commerce company illustrates this: during a rolling cache-cluster upgrade, outdated persistence files on nodes covering core shards caused hit rate to plummet from 99.9% to below 70%. Database CPU saturated, connection pools exhausted, and core interface timeouts spiked. The on-call engineer expanded the database three times before realizing the root cause was an empty cache, not database incapacity. Recovery took 40 minutes by shedding traffic and emergency backfilling hot data.

Postmortem revealed a key insight: cache data is regenerable, but regeneration takes time. During that window the system enters a dangerous state it never experiences normally. This article dissects the evolution of cache reconstruction: from cold-start endurance, to hot-data backfill, to dual-cluster hot standby that eliminates rebuild entirely, and finally to shard-level background rebuild — plus how to protect the database throughout the rebuild period.

What Exactly Is Cache Rebuild Reconstructing?

Unlike database recovery which targets RPO (how much data can be lost), cache recovery has no data-loss problem because every cache record can be re-derived from the database or compute layer. What must be rebuilt is the hit rate. At 10M QPS with 99.9% hit rate, the database sees only 10k QPS. If hit rate falls to 90%, the database faces 1M QPS (100×). At zero hit rate, the full 10M QPS hits the database (1000×). Database capacity is designed assuming a 99%+ cache hit rate; a cold cache removes that assumption.

Cold cache and empty cache differ: empty cache has no data; cold cache has data but the access structure isn't established (e.g., just-loaded persistence files full of unaccessed keys). Their danger levels and treatments differ.

The goal of cache rebuild is to move the system from "all requests hit the database" back to "most requests hit cache." Rebuild speed directly determines how long the database survives overload.

Why the First Few Minutes of Cold Start Are Fatal

Common cold-cache sources: node restart with stale persistence, cluster upgrade/migration starting from zero, traffic surge triggering mass eviction, dual-cluster switch to a lagging standby. The ensuing chain reaction is nearly identical:

Hit rate drops.

Hot keys are pierced first — thousands of concurrent requests for a missing key all backsource to the database.

Database overloads.

Upstream retries amplify load further (e.g., 10M QPS becomes 15M QPS).

Natural warm-up only works when backsource pressure stays within database capacity; at 10M QPS it almost never does. A hidden feedback loop worsens things: database slows → cache writes slow → hit-rate recovery curve flattens further. This "life-death window" is where every negative feedback accelerates.

Persistence Recovery: Looks Stable but Actually Slow

Redis RDB snapshots and AOF logs let a node reload its memory on restart. For small systems this suffices. At 10M QPS three problems emerge:

Load time: A 100GB instance takes minutes to load; cluster-wide simultaneous restarts contend for disk and network I/O, stretching each node's load time.

Data freshness: RDB is a point-in-time snapshot; writes between snapshots are missing. AOF can be finer but is larger and slower to replay. After load, stale data triggers another backsource wave.

Restores data, not access structure: Loaded memory may contain millions of cold keys while the few hot keys that actually carry traffic are a tiny fraction. If hot keys were evicted or the snapshot missed the latest hot distribution, hit rate stays low. The yardstick for cache recovery is the hit-rate curve, not memory usage; a cache full of cold data is as dangerous as an empty one.

Persistence recovery is a baseline for single-node failures but cannot handle cluster-level cold starts or protect traffic during rebuild.

Hot Data Backfill: Only Rebuild the Traffic-Bearing Portion

Turn passive warm-up into active backfill. Three sources for the hot-key list:

Access logs (offline frequency analysis over the last 24 hours).

Cache-internal hot statistics (e.g., LFU eviction metadata).

Business-explicit tags (e.g., promo item IDs pushed by operations before a big sale).

Internet access distributions are extremely skewed: 1% of keys often serve 80%+ of traffic. Rebuilding the full dataset is unnecessary; backfilling the hottest 1% can restore hit rate above 80%, leaving the long tail to natural warm-up.

This insight drives tiered backfill:

First batch: top 0.01% super-hot keys — plug the most dangerous penetration points within minutes.

Second batch: top 1% — bring hit rate to ~80%.

Subsequent batches: lower priority, yielding to online traffic.

Backfill itself must not kill the database. It is essentially a controlled backsource storm. If the database can spare 50k QPS, backfill rate is limited to 30–50% of that headroom, reserving the rest for live user backsource. Slower backfill is acceptable; killing the database loses everything.

Full flow: generate hot list → start tiered backfill → monitor hit rate and database load → keep traffic restricted (core interfaces only) until hit rate reaches a preset threshold (e.g., 95%) → gradually release traffic. At 10M QPS this typically finishes within 10 minutes.

Boundary: backfill depends on hot-list quality. If traffic distribution shifts due to the failure itself (e.g., a new promo with no history), the list becomes inaccurate. Mature systems bind backfill to business plans: predictable events have pre-prepared hot lists independent of historical stats.

Dual-Cluster Hot Standby: Replace Rebuild with a Switch

For payment and core trading paths requiring second-level recovery, the mindset shifts: don't rebuild after failure; keep a hot standby cluster always ready. Architecture: two cache clusters, primary serves all traffic, standby stays in sync via async event-stream replication. On primary failure, traffic switches in seconds — no cold start because the standby is already hot.

Cost objection: cache capacity doubles. But a single 10M QPS cache outage typically costs far more than a year of the extra cluster. For core paths the math is clear.

Replication leverages cache regenerability: a few seconds of lag is acceptable (worst case: one backsource or a stale read). Async event streaming keeps standby lag in seconds, imperceptible for most businesses.

Standby usage patterns: pure standby (idle, simple switch) vs. active-active (both serve traffic, mutual backup, higher utilization but more complex routing). 10M QPS systems favor active-active because pure standby clusters tend to get repurposed and become unavailable when actually needed — a recurring accident.

Boundary: dual-cluster solves cluster-level disasters (datacenter power loss, whole-cluster outage). Single-node/shard failures are handled by intra-cluster replicas (covered next).

Single-Shard Failure: Make Rebuild a Background Task

At 10M QPS, clusters have hundreds to thousands of nodes; single-node failure is daily routine. The system must auto-recover without impacting online traffic. Standard flow: detect → isolate → rebuild → verify. Three cache-specific traits:

Backsource handling during rebuild: Requests to the missing shard backsource to the database. Without control, a hot shard's absence opens a directed traffic hole. Request coalescing (singleflight) is mandatory: concurrent requests for the same key allow only one real database query; others wait for that result. This standard cache-penetration guard becomes critical during rebuild.

Rebuild traffic scheduling: One node hosts dozens of shards; a whole-node loss means dozens of simultaneous rebuilds. If all run at full speed, the database drowns in rebuild traffic alone. Solution: global rebuild queue, per-node and per-shard concurrency limits, dedicated rate-limit channel separate from online backsource budget.

Rebuild speed hard constraint: Must finish before the next failure on the same shard (otherwise service degradation becomes real even though data is regenerable). Target: minutes. Techniques: rebuild only hot data for that shard, multi-threaded parallel fetch, pull baseline from a nearby replica instead of the database.

The highest form of failure recovery is not strong emergency response, but making recovery a routine background mechanism so failures never require emergency handling.

How to Protect the Database During Rebuild

Rebuild takes time; the database must survive that window. Four protective layers form a "speed bump" not a "firewall": they cannot prevent slowdown but prevent death.

Backsource rate limiting: Define a database backsource budget (e.g., 120% of normal load). Excess requests are intercepted. Degradation options: return empty/default, return stale value (most common), or error for retry.

Serve stale (expired-but-usable): Relax expiration logic — continue serving expired values, queue or drop refresh requests. Keep an "expired but usable" copy in cache or app-local memory. Data a few seconds/minutes old is far better than total outage.

Request coalescing (singleflight): As above, collapses thousands of concurrent hot-key requests into one real backsource, directly eliminating database amplification.

Tiered fallback: App-local cache (seconds), static pages/CDN (block read traffic before it reaches app), degrade non-core features to reserve resources for core paths.

The bottom line of cache rebuild is not how fast it rebuilds, but that the database must not die; if the database lives, hit rate recovery is only a matter of time. These protections also let rebuild tasks run more conservatively (slower backfill, better ordering) without rushing to extremes. Protection and rebuild mutually reinforce each other.

Make Pre-Warming a Routine Capability

The toolkit is now complete: persistence for single-node, hot backfill for cluster cold start, dual-cluster for major disasters, shard-level background tasks for daily wear. Ensuring they work when needed requires drills and metrics.

Hit rate is the core SLO for cache recovery. Mature teams define a recovery target (like RTO): minutes to restore hit rate >95%. This number comes from drills: regularly clear a shard and observe the recovery curve, simulate full cold start on a shadow cluster, deliberately empty cache during full-chain stress tests before big promotions. Record actual recovery time and compare with previous runs.

Pre-warming also serves planned changes: cluster scaling (new nodes are cold), shrinking, migration, version upgrades — all involve data movement and should reuse the same pre-warm system. Planned pre-warm and unplanned rebuild use the same system capability; only the trigger differs.

Across 100k, 1M, and 10M QPS stages, the evolution is consistent: recovery executor shifts from human to system, timing shifts from post-incident to pre-incident, granularity shifts from whole cluster to single shard. The endpoint of pre-warming is eliminating the need for pre-warming; hot standby and routine rebuild make "cold cache" virtually disappear in production.

Keep the Cache Forever Hot

Revisiting the opening incident: every layer could have caught it — verifying persistence freshness before upgrade, hot backfill to recover in 10 minutes, dual-cluster to avoid restart entirely. The company had technical capability but lacked a systematic cache-recovery construction; mechanisms were scattered across teams, not linked into a complete recovery chain.

For 10M QPS systems, cache is the database's shadow guard; when the guard falls, the master is defenseless. Cache-recovery metrics (hit-rate recovery time, backfill bandwidth, hot-standby switch time) should enter architecture reviews as hard criteria, backed by drill data.

If your system only has persistence recovery, start with a real cold-start drill: clear a non-core shard, observe the hit-rate curve and database load, measure the true recovery time. That number will likely force a reprioritization of cache-failure readiness. When was your cache cluster last fully cleared?

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

system architecturecold starthigh QPSsingleflightcache warmingdatabase protectioncache reconstructiondual-cluster hot standbyhot data backfillshard-level rebuild
Random Bulletin
Written by

Random Bulletin

17-year internet software developer specializing in AI applications, networking, architecture, and open source. Led the delivery of network services handling hundreds of millions of concurrent devices and tens of millions of QPS, and has three years of experience designing and building an agent platform. Follow to stay updated.

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.