From Static DNS to Dynamic Traffic Scheduling: Mastering 10M QPS Failover
This article dissects why static DNS-based traffic scheduling fails at ten-million-QPS scale and how layered dynamic scheduling — across access, gateway, service, and data layers — combined with unitized routing, automated decision loops, and rigorous drill practices enables precise, reversible, and safe failover.
The article opens with a real同城 (same-city) disaster recovery drill at 2 AM: database primary‑secondary switch completed, capacity reserved, yet after ten minutes only 30 % of expected QPS reached the standby site while 60 % still hit the logically offline primary, error rates climbing. Root cause: DNS weighted resolution with 600 s TTL, recursive caches ignoring TTL, client‑side connection reuse — a long‑tail propagation curve where each minute at 10 M QPS means hundreds of millions of failed requests.
Disaster Recovery Loop: Scheduling Is the “Hand”
A full DR response is a closed loop: detection → decision → scheduling → verification. The loop’s total latency equals its slowest link, and scheduling is the most underestimated. Many teams over‑invest in redundant capacity and data sync but rely on runbooks and DNS records for traffic movement — resulting in standby capacity sitting idle while traffic cannot shift.
Static scheduling is “draw a map and let traffic find its way”; dynamic scheduling is “place traffic police at every intersection to redirect flow in real time.”
Three Fatal Flaws of Static Scheduling
Propagation latency: DNS chain (client → LocalDNS → authoritative) has caches at every hop; TTL is advisory. Long‑lived connections never re‑resolve. Real generation curve: ~60 % at TTL mark, remainder minutes to tens of minutes.
Unobservable, uncontrollable: DNS change is broadcast; no visibility into which clients have switched, no way to shift only 5 % first. Rollback suffers same latency, doubling the fault window.
Coarse granularity: DNS only distinguishes geography (via LocalDNS) and weight. Cannot express “cut order traffic but keep payment”, “isolate a large tenant”, “cut read only”. Real stop‑the‑bleed scenarios demand fine‑grained traffic surgery.
Control leakage: authoritative server updates, but recursive caches belong to ISPs with arbitrary implementations. The slowest, least controllable link dictates actual cutover speed — an architectural original sin of static scheduling.
Note: moving the switch in‑house means you must guarantee the switch itself never jams — control‑plane failure is the highest‑severity accident in dynamic scheduling.
Layered Dynamic Traffic Scheduling
Core principle: every layer the request traverses must be able to change its destination by rule. Four layers, each with distinct mechanisms, granularity, and speed:
Access layer — “which site”: dynamic GSLB (real‑time probe‑based resolution), BGP Anycast (same IP multi‑site, withdraw route on failure, network‑layer convergence in seconds), self‑built health checks. Covers all ingress including long‑connections and SDK direct links. Anycast is fastest site‑level fallback.
Gateway layer — “which cluster/group”: full request context (headers, URL, tenant, region tags) as routing basis; config‑center push, seconds to full cluster. Used for fine‑grained stop‑the‑bleed: route away from unhealthy downstream clusters.
Service layer — “which instance”: RPC routing rules, load‑balancing policies, swimlane isolation; per‑instance, per‑interface, millisecond‑to‑second effect. Daily use: canary, A/B, evict bad instances.
Data layer — “who owns write rights”: master‑secondary switch, position catch‑up, split‑brain protection; minute‑level, must be pre‑planned. Data scheduling is not traffic scheduling — it is the boundary traffic scheduling must obey: traffic can go anywhere with capacity, but writes only where data ownership allows.
Access layer deep dive: two mainstream paths combined in mature systems.
GSLB path: self‑operated global LB, distributed probes measure site access quality, resolution dynamically generated per region/ISP/capacity. Strong control, fine‑grained; cost: build/maintain probe+resolution stack; still bound by DNS cache TTL (10 s–minutes).
Anycast path: same IP announced via BGP at multiple sites; network routes to “nearest”; withdraw route on failure, convergence seconds to tens of seconds, zero DNS cache dependency. Granularity limited by network topology; withdrawal may cause detours to distant sites, raising latency.
Best practice: Anycast for instant site‑level safety net; GSLB for precise steady‑state placement. Both rely on a multi‑point health probe system covering major ISPs/regions, capable of distinguishing real site failure from probe‑path issues.
No single layer suffices: only access layer → cannot shift fine‑grained traffic; only service layer → site‑level disaster blocks ingress; only data layer → traffic never reaches correct destination.
Unitized Architecture: Turning Traffic into “Containers”
Layered scheduling solves “can we shift”; unitization solves “what unit to shift”. A cell (unit) is a self‑contained deployment unit with closed traffic, service, and data loops. From scheduling view, it turns discrete traffic into standard containers that can be moved whole.
Without unitization, scheduling operates on scattered dimensions (domain, interface, instance); a site cutover becomes a full‑site rewiring. With unitization, scheduling unit becomes the cell: internal closure means no cross‑cell dependencies; scheduler only answers “which site hosts this cell”.
Shard key (e.g., user_id) drives two‑level mapping: shard key → cell → site. Routing rule table is the core asset. Steady state stable; failure = rewrite table: map cells from failed site to healthy site. One table change + one rule push = whole traffic migration, no per‑service verification.
Two often‑confused concepts:
Traffic coloring: tag request with lane/gray‑batch/drill ID so every hop identifies and handles consistently. Coloring = identification; routing rules = destination. Both needed for “send specific batch precisely to specific place”. Drill traffic must be fully colored to prevent leakage.
Global vs. unit services: Global services (auth, config, risk lists) cannot be containerized; they use multi‑site active‑active + nearest access. Unit services migrate via rule table rewrite. Mixing their scheduling logic is a classic failure source.
Routing rule table must be governed like DDL: versioned, auditable (who/when/why), one‑click rollback to any version, shadow‑environment simulation before push, fallback pull channel if config center fails. Every scheduling component must answer: when you break, can traffic scheduling still happen? — the recursive hardness unique to dynamic scheduling.
Unitization transforms “whether to cut” into “which cells, how much each” — a multiple‑choice question that can be taken in small steps, verified, and rolled back.
Decision Loop: From Manual Approval to Automated Execution
Early model: human watches monitor, assembles war room, confirms layer by layer, manually executes — 10–30 min decision chain, while 10 M QPS losses accrue per second. Full automation risks false positives: moving traffic from healthy sites, causing “self‑healing” more deadly than the fault.
Mature four‑level split by risk:
Instance level — fully auto: LB health check evicts bad instance in seconds; misjudgment cost = one instance capacity; auto‑retry safety net.
Cluster level — auto + approval: system proposes rule change, human one‑click confirm; misjudgment cost grows.
Unit level — plan‑driven: each unit’s migration path, capacity check, data check pre‑written as plan; one‑click execute with pause/rollback mid‑flight.
Site level — human decision, auto execution: impact too large, often involves data layer; human judges, but execution is platform‑driven one‑click to avoid manual errors.
Three engineering guards for credible automation:
Redundant criteria: single probe source untrustworthy; auto‑trigger requires consensus from multiple independent signals (intra‑DC probe, cross‑DC probe, ISP dial‑test).
Switch budget: rate‑limit auto migrations (e.g., max one per unit per hour) to prevent thrashing at boundary conditions.
Canary validation: even after trigger, shift 5 % first, observe core metrics, confirm target site truly absorbs, then ramp.
Asymmetric misjudgment cost: miss (should cut, didn’t) = linear business loss over time; false cut (shouldn’t cut, did) = instant shock + potential secondary disasters, often amplified. Hence automation leans conservative: higher thresholds, longer confirmation windows, “slower cut” for “almost never false cut”. True speed needs delegated to low‑impact instance/cluster auto.
Post‑cut verification (often skipped): confirm (1) target site core metrics (success rate, latency, error rate) stable under new load; (2) zero residual traffic on old path; (3) no unexpected cross‑site calls from routing change. All three pass = cut complete; otherwise “in‑progress action”.
Platformization: a dedicated scheduling platform structures all plans (trigger conditions, steps, capacity checks, rollback), white‑boards execution (step status, effective nodes, metric curves, pause/rollback buttons), archives decision trail (who/when/basis/plan). Turns cutover from hero‑dependent improvisation into system‑backed standard motion.
Secondary Disasters: Traffic Tides & Stampedes
Once scheduling works, the enemy shifts from “can’t cut” to “cut too fast”. Four classes of self‑inflicted disasters:
Traffic tide: receiving site instantly doubles load. If capacity designed at 1.5×, the 0.5× gap breaches in minutes: CPU saturation, cache hit‑rate collapse, connection pool exhaustion, cascading failure. Mitigation: batched migration + pre‑flight capacity check — scheduler must answer “current water level, post‑cut water level, what if over red line” — refuse or partial cut if over.
Cold‑start stampede: target instances long idle: cold JVM, empty cache, unbuilt connection pools. Real throughput far below design. Mitigation: pre‑warm + slow ramp — inject trickle traffic pre‑cut to load cache, JIT compile, establish connections; then stair‑step: 5 % → observe 10 min → 20 % → 50 % → 100 %, each gate gated by core metrics.
Control‑plane split‑brain: multi‑site schedulers each think they’re master, push conflicting rules. Fix: control plane cross‑site deployed with clear leader election; extreme case (control plane totally down) → data plane runs on locally cached rules, correct posture on control‑plane loss is “freeze status quo”, not “each decides independently” .
Oscillation & flip‑flop: probe signals jitter at threshold; without hysteresis, traffic ping‑pongs, each move costs and risks. Fix: hysteresis bands + cooldown — e.g., evacuate at >80 % water level, allow return only <60 %, mandatory cooldown between moves.
Maturity of a scheduling system is not measured by how fast it cuts, but by whether it dares to cut in small steps, can pause anytime, and can one‑click rollback. Speed is outcome; controllability is capability.
Drills: Turning Plans into Muscle Memory
Rule tables with hundreds of mappings, plans with dozens of branches, auto policies with ten-plus triggers — hidden errors only drills expose. Three drill tiers:
Tabletop: scenario on paper, on‑call walks decision flow; tests plan readability & decision smoothness; lowest cost, high frequency.
Real traffic cut drill: production environment, shift real traffic subset; validates actual scheduler effect speed, target site real absorption, monitoring linkage; must run off‑peak, full traffic coloring, instant rollback ready.
Surprise drill: no advance time notice; dedicated team injects fault; measures end‑to‑end real reaction time from detection → decision → scheduling execution, plus on‑call response capability.
Quantified metrics (must be tracked separately to locate bottleneck):
Decision latency (fault confirmed → command issued)
Rule full‑effect latency (command → all data planes ack)
Traffic migration latency (rule effective → traffic distribution hits target ratio)
Business loss integral (error‑rate area + success‑rate dip duration)
These metrics feed back into capacity planning: if target site water level nears red line → capacity insufficient, not scheduling issue; if rule push fast but traffic migration slow → bottleneck in access‑layer generation path, not decision layer. Drill metrics turn “DR capability” from vague adjective into numbers that enter architecture reviews and year‑over‑year comparisons.
Drills also continuously calibrate the team’s “cut confidence”. A team that cuts real traffic quarterly hesitates orders of magnitude less during real faults than a team that only saw cutover in docs.
Evolution Ladder: From DNS to Unitized Scheduling
Each step trades complexity for certainty and granularity:
GSLB first — highest ROI; no business architecture change; dynamic DNS alone compresses site‑level cut from uncontrolled long‑tail to minutes.
Gateway & service rule‑based scheduling — enables fine‑grained stop‑the‑bleed; depends on config‑center push + full‑chain tagging system.
Unitized scheduling — most expensive; requires prior data ownership clarification and unitization refactor; scheduling merely turns refactor results into executable capability.
Cold water: if still monolithic, data unsharded, services mesh‑dependent, unitized scheduling is a castle in the air. Scheduling ceiling is set by how regular the scheduled object is. First make traffic into movable containers, then build the gantry crane — order cannot reverse.
Scheduling limit = capacity floor; capacity plan must sit at same table as cutover plan. Common mismatch: capacity team designs per‑site standalone full‑load redundancy; scheduling team designs “cut half traffic anytime” plans; combined they deadlock because redundancy only 1.5×.
Case Study: One Team’s Journey
Post‑mortem of opening drill led to three‑step improvement:
Access layer: static DNS → dynamic GSLB + Anycast fallback; site‑level cut generation from uncontrolled long‑tail to <2 min.
Gateway layer: rule‑based routing (by ratio, by tenant); cutover from all‑or‑nothing broadcast to batched verifiable action.
Unitization: scheduling unit collapsed to cell; 10‑page runbook → one‑click on scheduling platform.
One year later, same site‑level scenario: decision‑to‑migration complete in 8 min, with a 5 % canary validation and a deliberate pause mid‑way.
Traffic scheduling’s journey from static to dynamic appears as a tech‑stack upgrade; fundamentally it is DR shifting from “relying on assets” to “relying on capabilities”. Redundant capacity is an asset lying in the data center — it doesn’t self‑convert to safety. Scheduling capability is the hand that cashes the asset into insurance. How steady, fast, and controllable that hand is decides whether your DR system is a real policy or a pretty diagram.
Ask your system today: a site‑level cut — how long, how many approval layers, how many manual steps? Those numbers are the most honest measure of your current scheduling capability.
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
Random Bulletin
17-year internet software developer specializing in AI applications, networking, architecture, and open source. Led the delivery of network services handling hundreds of millions of concurrent devices and tens of millions of QPS, and has three years of experience designing and building an agent platform. Follow to stay updated.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
