Service Degradation at 10M QPS: From All-or-Nothing to Partial Availability
At ten million QPS, a single slow dependency can cascade into full outage; this article details a systematic degradation framework: business capability grading (L0–L3), degradation contracts with metadata, budget-constrained action ladders, automated controllers with leases and hysteresis, multi-level recovery state machines, cross-service guardrails, and validation via chaos drills.
Why "All-or-Nothing" Amplifies Small Failures
A product detail page aggregates eight downstream services: product, price, inventory, marketing, reviews, recommendations, membership, and risk control. Normally they respond in ~50 ms, but when the recommendation service degrades to a 3 second P99, the aggregator blocks on the slowest call. Threads, connections, and memory are held for the full duration. Using Little's Law: at 1 million requests/sec, average latency rising from 100 ms to 2 seconds inflates in-flight requests from ~100 k to ~2 million, exhausting connection pools and starving healthy downstream calls.
Retries compound the problem. Client, gateway, and framework retries multiply traffic; the failing dependency never recovers, and the entire chain processes doomed requests. The core insight: a complete response ≠ core value. Users need buyable status, price, and inventory; personalized recommendations are enhancements, not prerequisites.
Grade Capabilities Before Defining Degradation Switches
Four Levels Beat Binary "Core/Non-Core"
Levels are defined by business outcome and degradation mode:
L0 – Must Succeed : payment, inventory deduction, core risk control. Zero tolerance for degradation.
L1 – Degraded Function : pricing, promotions, core search. Can fall back to conservative rules or cached results.
L2 – Simplified Experience : personalized recommendations, rich content. Can serve stale data, rule-based fallbacks, or hide the module.
L3 – Optional Enhancement : analytics, logging, non-critical decorations. Can be disabled entirely.
Grading is contextual: the same "promotion" capability is L1 at checkout but L2 on a history page. The degradation unit is scenario × capability × tenant/region , not the whole service.
Every Capability Needs a Testable Degraded Result
"Ignore the exception" is insufficient. The aggregator must know what to return, the frontend must handle missing fields, and product must accept the UX. Common degraded results:
Last successful result with a timestamp ( data_age)
Rule-based or global hot rankings
Truncated lists (top‑N only)
Disable personalization, reuse public cache
Conservative static rules instead of real-time computation
Hide the module but keep page skeleton
Queue writes or switch to read-only
Degraded results must be independently testable; a frontend that refreshes on empty arrays or renders missing prices as ¥0 defeats the purpose.
Degradation Contracts in Interface Semantics
Callers must distinguish three outcomes: normal, acceptable degraded, and hard failure. Response metadata carries degraded, reason, data_age, policy_version so every layer knows the reliability of the payload. Only when business has defined "what partial availability looks like" can the technical system automatically enter partial availability.
Degradation Is a Ladder of Actions Ordered by Loss
Reduce Cost Before Dropping Capability
Each capability has internal cost tiers: search can reduce recall channels, rerank depth, or result count; recommendation can drop real-time features; reports can switch to hourly snapshots. These preserve partial value while releasing CPU, memory, downstream QPS, or bandwidth.
Different dependencies demand different strategies (illustrated in the article's diagram). "Fail-open" looks like high availability but can turn technical faults into financial or compliance risks. The goal is not maximal success rate but an acceptable trade-off among business loss, data correctness, and system capacity.
Degradation Budgets Constrain Actions
Each scenario defines a budget with four dimensions:
Maximum affected traffic percentage
Maximum duration
Acceptable data staleness
Non-negotiable business/compliance boundaries
Actions exceeding the budget escalate to human approval or a more conservative strategy. Automation handles only low-risk, repeatable decisions.
Explicit Action Order
Under pressure, execute in sequence:
Reduce log sampling, debug info, non-essential internal computation
Throttle or disable L3 capabilities
Switch L2 to cache, defaults, or static results
Restrict low-priority tenants and non-core entry points
Simplify L1 high-cost paths, enable conservative rules
Admit-control L0 to prevent total overload
Each step observes whether resources return to a safe zone before proceeding.
How the Automatic Controller Decides "What to Reduce Now"
Trigger Signals Must Cover Demand, Supply, and Outcome
Demand : ingress QPS, concurrency, request cost, tenant distribution.
Supply : healthy instances, CPU, connection pools, queues, downstream quotas.
Outcome : core success rate, tail latency, rejection rate.
Example: high CPU but stable core success rate → pause offline jobs instead of user features. Conversely, 55% CPU but connection pool saturated by a slow dependency → fast-fail and serve cached results.
Actions Declare Their Resource Yield
Each degradation action publishes expected gains: downstream QPS reduction, freed concurrency slots/connections, CPU/memory savings, cross-region bandwidth reduction, and impact on core success rate/revenue. The controller models selection as a constrained optimization: minimize business loss while regaining sufficient headroom. Production need not achieve mathematical optimum; ranking by estimated yield and correcting from observed effect is far superior to ad-hoc switch flipping.
Decisions Carry Leases, Not Perpetual Effect
Every automatic degradation records trigger cause, target scope, policy version, start time, and expiry. On expiry the controller re-evaluates; manual extensions require justification. Multiple controllers (capacity, circuit-breaker, event-protection) may act simultaneously; a single target state or explicit priority merge rules prevent oscillation.
Shadow-Run Before Auto-Execute
New policies run in shadow mode: compute actions but do not execute. Compare suggestions against real incident handling to measure false-positive rate, yield accuracy, and blast radius. Then enable only low-risk, small-blast-radius actions, gradually expanding automation. The degradation platform itself is a high-risk control plane and deserves its own canary rollout.
Make Degradation a Recoverable State Machine
Enter Fast, Exit Slow
Entry uses a short window (e.g., three consecutive 10-second windows breaching threshold). Exit requires a longer stable window (e.g., five continuous minutes in the safe zone). Separate entry/exit thresholds create hysteresis, avoiding flapping at the boundary.
Recovery Is Staged Ramp-Up
Do not restore all functions at once. Ramp: 1% → 5% → 20% → 50% → 100% of traffic on the full path, with a minimum observation window at each step. Roll back to the previous stage on anomaly.
Handle Cache and Backlog Explicitly
Stale caches, async queues, and unfinished tasks accumulated during degradation do not vanish. Before recovery, decide:
Let old caches expire naturally or actively refresh?
Refresh with rate limiting and random jitter?
Process backlog in order or discard stale tasks?
Do delayed writes still have business meaning?
How to allocate capacity between live traffic and compensation tasks?
Reserve fixed capacity for online traffic; limit compensation to the remaining headroom. Degradation is complete only when full capability is restored and core metrics remain stable.
At 10M QPS, Degradation Is No Longer a Single-Service Problem
Same Action, Different Yield Across Failure Domains
If one data center's recommendation service fails, degrade only that DC's traffic; global disable wastes healthy regions. A tenant's spike should not force all users into simplified mode. Policy scope must combine region, AZ, version, tenant, entry point, and request type. Fine granularity explodes policy count; the platform needs inheritance/override rules (global default → region override → VIP exception → incident-time forced policy) with explicit priority to avoid unpredictable conflicts.
Degradation Actions Shift Pressure Across Services
Disabling real-time recommendation may increase public cache reads; switching reports to snapshots may spike object-storage bandwidth; async queues may accumulate. Local CPU relief ≠ system health. The controller must observe cross-service resource changes post-action to avoid merely moving the bottleneck.
Multi-Level Degradation Needs a Global Resource Ledger
At scale, healthy capacity is simultaneously consumed by failover traffic, canary releases, promotional spikes, and compensation tasks. A single service sees only its own CPU, unaware that another control plane is about to shift 30% more traffic its way. A global ledger aggregates per-fault-domain: available capacity, committed capacity, active actions, and estimated recovery time.
Guardrails Define How Far Automatic Degradation Can Go
Hard Boundaries That Cannot Be Crossed
L0 capabilities cannot be disabled by ordinary automatic policies.
Price, funds, permissions fields forbidden from using sourceless defaults.
Single action traffic impact ≤ configured ceiling.
Concurrent degradation levels per scenario ≤ ceiling.
Stale or missing metrics → prohibit action expansion.
Insufficient healthy capacity → prohibit shifting more traffic into the domain.
Every policy must have a lease, audit trail, and explicit owner.
The execution layer also provides an emergency stop: freeze new automatic actions while retaining already-effective, still-protective measures. Blindly reverting all switches may remove active mitigations.
Degraded Results Must Meet Correctness Floors
Available ≠ correct. Stale inventory cache → oversell. Fail-open risk control → financial loss. Permission service default-allow → security breach. Each policy defines a correctness boundary; beyond data freshness or risk threshold, reject the request rather than return misleading results.
Conservative Mode When Metrics Lie
Automation depends on monitoring. If metrics are delayed or the observability pipeline is overloaded, the controller decides on a stale world. Policy inputs carry collection timestamp, window, and completeness. On stale data: maintain existing leases, pause action expansion, alert humans.
Verifying That Degradation Actually Reduces Loss
Observability Across Four Layers
(Diagram in article shows four layers: infrastructure, platform, application, business.) Focusing only on error rate misleads. Degradation may raise HTTP success rate while serving empty pages, or deliberately reject low-priority traffic making aggregate error rate look worse while core transactions recover. Metrics must be interpreted against business objectives.
Drills Must Cover Both Failure and Recovery
Inject progressively:
Single weak dependency latency spike → verify fast main-path return.
Cache hit-rate drop → verify fallback doesn't reverse-penetrate storage.
Fault-domain capacity loss → verify policy scope precision.
Metric delay/loss → verify controller halts expansion.
Concentrated cache refresh on recovery → verify rate limiting and jitter.
Two policies match simultaneously → verify priority merge and target state.
Record not just "pass/fail" but trigger-to-effect latency, actual vs. estimated resource release, business loss, recovery duration, and pressure migration. Data feeds back to calibrate action yields.
Phased Rollout: From Manual Switches to Closed Loop
Phase 1 : Eliminate "no fallback" dependencies; ensure every call returns a defined result on failure.
Phase 2 : Consolidate scattered switches into a unified platform with RBAC, audit, and TTL.
Phase 3 : With sufficient data, automate low-blast-radius, easily-rolled-back actions. Controller permissions grow with evidence, not with feature launch.
From Complete Features to Stable Outcomes
"All features" is ideal in steady state, but fault handling optimizes for preserving the most important business outcomes, not maintaining a feature checklist. Under resource pressure the system can do less, update slower, return simpler data — while avoiding total collapse.
The foundation is not a switch platform but business grading, degradation contracts, and clear correctness boundaries. Automation adds evidence-based judgment, resource-yield estimation, leases, guardrails, and recovery state machines. As scale grows, the system must also handle fault-domain variance, cross-service pressure migration, and control-plane conflicts.
The real post-mortem question: If the slowest dependency fails again tomorrow, can your core path accept "partial" within seconds, instead of losing "all" together?
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
Random Bulletin
17-year internet software developer specializing in AI applications, networking, architecture, and open source. Led the delivery of network services handling hundreds of millions of concurrent devices and tens of millions of QPS, and has three years of experience designing and building an agent platform. Follow to stay updated.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
