Fault Prediction at 10M QPS: From Zero to Production-Ready System
This article details a practical roadmap for building production-grade fault prediction systems at massive scale, covering target selection, data governance, model evolution, time-window design, and safe action loops—emphasizing that reliable prediction requires engineering rigor beyond just model training.
The article opens with a familiar scenario: a 2 a.m. disk-latency alert cascades into a full-chain outage within minutes, prompting the postmortem question—why wasn’t the rising trend caught earlier? The author reframes the problem: fault prediction is not about guessing the next accident, but about providing a credible, explainable, and actionable evidence chain while risk is still controllable.
What Prediction Does Beyond Alerting
Alerting answers “is it abnormal now?”; prediction answers “is the risk of a specific fault class rising in the next N minutes?” The same signals (CPU, queue length, GC pauses) can support both, but the judgment target differs. A CPU >90% alert judges current state; a rising baseline over two hours with growing run-queue and GC pauses—without matching traffic growth—supports a “capacity-type fault in 30 minutes” judgment. Prediction also differs from capacity planning: planning uses business growth and budgets for weeks/months; prediction uses live telemetry and recent changes for minutes/hours. The three layers—thresholds, anomaly detection, prediction—are complementary, not replacements.
Selecting Faults Worth Predicting
Teams often start by collecting all data and hunting for the best model. The author argues the correct starting point is a fault catalog listing: object, failure mode, precursors, achievable lead time, and executable actions. A fault is suitable for prediction only if:
It has an observable gradual process (disk bad sectors, memory leaks, cert expiry, thread-pool backlog, connection-quality decay). Sudden power loss or fiber cuts lack a stable prediction window.
Knowing earlier creates action space. A 10-second heads-up is useless if human response takes five minutes; 20 minutes enables shard migration, batch-job throttling, or standby warm-up.
False-positive and false-negative costs are calculable. Auto-evicting a healthy node is riskier than paging a human; the same model needs different thresholds for each action chain.
The author introduces a “prediction contract” per fault class. Example for slow storage disks: object = physical disk; fault event = sustained high tail latency for 10 minutes causing replica catch-up lag; prediction window = next 30 minutes; minimum lead time = 15 minutes; high-risk results go to human review first, no auto-eviction. This turns vague “early detection” into a verifiable engineering target.
Why Data Is Harder Than Models
Stable Entity Identity
A single machine may have hostname, instance ID, container node name, asset tag, and cloud resource ID. Services scale, pods restart, disks remount. Without a stable entity-resolution layer mapping telemetry, config, changes, topology, and events to a canonical resource identity, training data stitches unrelated histories or splits one object into fragments. Short-lived instances must retain context: service, version, hardware type, failure domain. Prediction targets may be shards, clusters, services, or failure domains—not just single instances.
Time Alignment by Event Time, Not Ingestion Time
Metrics, logs, traces, and change events arrive with different latencies (logs buffered for minutes, config events second-granularity, cross-region clock skew). Joining features by ingestion time leaks future information into training. Production features must satisfy “what was truly available at prediction time.” This requires storing both event time and ingestion time, using watermarks for late data, and ensuring offline training and online inference share identical window definitions.
Positive Labels Cannot Rely Solely on Incident Tickets
Tickets describe impact and resolution, not precise resource and start time. A service outage caused by a single disk jitter may be labeled only “API timeout.” Labeling all nodes during the incident window injects massive noise. Usable labels come from multi-source evidence: alert states, auto-recovery actions, resource replacement records, error-budget burn, human confirmation, postmortem conclusions. Labels must distinguish root-cause objects from affected objects, otherwise the model learns propagation symptoms.
Negative Sampling Requires Care
Treating all “no incident” periods as negatives creates extreme imbalance and may include unrecorded micro-faults. Common practice: stratified sampling by object and time, preserving representative intervals across load, hardware, version, and region; down-weight or exclude windows with uncertain labels.
Time Windows Determine Prediction Realism
The most common mistake is random train/test splits. Adjacent windows are highly correlated; random splits place the same fault trajectory on both sides, inflating offline metrics while the model fails in production. Correct evaluation uses temporal splits: earlier data for training, later for validation, newest for test. If architecture upgrades, hardware swaps, or collection changes occurred, group tests by version and failure domain to detect overfitting to a specific cohort.
Each training sample involves three windows:
Observation window: history visible to the model (e.g., past 2 hours).
Gap window: safety buffer between prediction time and risk-window start, preventing the model from using already-manifested symptoms as “prediction.”
Prediction window: future interval the model judges (e.g., next 30 minutes).
Without a gap window, the model may rely on “error rate already spiked,” merely renaming detection. Too long a gap weakens visible precursors. Windows are not fixed hyperparameters; they must be validated against action latency and fault evolution speed.
Accuracy is misleading due to class imbalance. Meaningful metrics include: precision@k, recall@k, lead-time distribution, and calibration curves. Window-level recall is inflated by long faults (ten consecutive windows hit = one event). Production evaluation merges adjacent risk windows into events, measuring whether each event was caught early and the earliest lead time.
Evolving from Rules to Models
Layer 1: Interpretable Rules
Rules suit clear precursors with stable ops experience: disk reallocated sectors rising with I/O tail latency; cert expiry below threshold; memory monotonic growth under steady load. Drawbacks: high maintenance, rigid combinations. First version should decompose signals into five evidence categories—trend, volatility, saturation, errors, changes—each outputting a 0–1 score, then combine with transparent weights. This preserves explainability and reveals which evidence class degrades over time.
Layer 2: Statistical Baselines
Same 70% CPU may be normal for batch nodes but anomalous for a control plane usually at 15%. Rolling quantiles, seasonal baselines, exponential smoothing, and change-point detection upgrade fixed thresholds to “relative to own history.” Statistical methods need few fault labels, good for cold start. But anomaly ≠ fault; releases, promotions, batch jobs cause legitimate deviations. Statistical scores must be combined with change calendars, business activity, and topology impact—never trigger high-risk actions alone.
Layer 3: Supervised Learning
With stable labels, logistic regression or gradient-boosted trees learn multi-signal combinations. For tabular ops data, simple models often suffice; training, interpretation, and online inference costs are low. Inputs: window statistics, rates of change, year-over-year/week-over-week ratios, error proportions, dependency health, recent changes. Deep sequence models fit high-signal, complex-pattern, large-sample scenarios but are not default. They are more sensitive to collection drift and harder to explain. If simple models show no bottleneck on event-level metrics, added complexity rarely yields operational gains.
Model complexity should be driven by real bottlenecks in false positives, false negatives, and lead time—not by algorithm novelty.
Anatomy of a Production Prediction Pipeline
The pipeline has six independently observable stages: data ingestion, entity resolution, feature computation, risk inference, event aggregation, policy execution. The prediction service itself can fail.
Ingestion retains raw events for replay when feature logic changes.
Entity resolution maintains resource identity and topology snapshots.
Feature layer serves both offline and online, reusing computation definitions to avoid train/serve skew.
Inference outputs risk score, prediction window, model version, and top evidence—does not decide actions.
Event aggregation is critical: per-minute scores on the same node would cause alert storms. Aggregation deduplicates, suppresses, manages state transitions (observing → elevated → high-risk → mitigated → recovered), and applies hysteresis to prevent threshold flapping. For interdependent nodes, topology-aware merging prevents one root cause from generating dozens of predictions.
Policy engine decides based on risk, confidence, blast radius, and action cost. Key principle: separate “model score” from “action authority.” Model states risk; policy decides log, notify, create ticket, trigger drill, or execute auto-migration.
Safe Degradation of the Prediction Service
When feature latency spikes, model registry unavailable, or entity mapping breaks, the pipeline must not block business nor execute stale scores. Degradation: stop new automated actions, retain threshold alerts and detection, explicitly mark prediction data stale. Monitor not only inference latency but feature freshness, entity mapping coverage, model load success rate, score distribution, risk event count, and action success rate. A metric stuck constant may indicate broken collection, not stability.
Why 10M QPS Demands More Than Per-Node Scores
At 100k instances evaluated every minute, even 0.01% per-object false positive rate yields thousands of daily risk windows—overwhelming on-call. Three mechanisms are essential:
Edge/node-side low-cost features (sliding means, quantiles, rates of change) reduce high-cardinality raw data transport.
Service/shard/failure-domain aggregation distinguishes isolated node issues from common-upstream problems.
Heavy models and topology analysis run only on candidate risk objects.
Global action budgets limit high-risk events, migrations, or scale-outs per time window. Otherwise, widespread telemetry drift could trigger mass operations on healthy resources, creating new capacity oscillations. Correlation awareness: when a region’s network metrics degrade together, per-node models may flag thousands of machines. Topology aggregation should emit a single region-level event and pause cross-region auto-migrations that would worsen traffic shift. At scale, model drift on a specific hardware type, version, or region can be masked by global averages. Monitoring must slice by key dimensions—comparing recall, precision, and calibration per slice—to prevent the model from only working on the majority cohort.
Moving from 1M to 10M QPS isn’t just scoring ten times more objects; it requires upgrading from per-object prediction to a system jointly constrained by failure domains, topology, and action budgets.
Turning Risk Scores into Reliable Actions
Value materializes not in a dashboard curve but in reduced user impact or manual toil. Actionability must start low-risk.
Shadow run: model computes risk but takes no action. Align predictions with real events; validate lead time, false positives, data gaps, evidence explanations. Shadow period must cover normal days, release days, traffic peaks, and several real faults—not just a week of average metrics.
Assisted decision: high-risk events appear on on-call panel with object, window, top evidence, similar historical cases, suggested checks. On-call confirms or rejects, recording reason. Feedback is not blindly relabeled; distinguish model error, data error, fault preempted by action, and acceptable business variance.
Limited automation of reversible, low-blast-radius actions: add read replica, throttle non-critical batch jobs, pre-warm a few standbys. Each action has scope, rate, duration, rollback conditions, full audit trail.
High-impact actions (primary shard migration, node eviction, cross-region failover) only after stable model performance, gated by policy engine, approval, or dual-signal confirmation. Prediction scores never wire directly to infrastructure control plane.
Action effectiveness must enter evaluation. A high-risk prediction followed by no fault could be a false positive—or a successful pre-scale. Without counterfactuals, the system mislabels successful interventions as errors. Retain untreated control groups or use phased rollouts within safety bounds to estimate whether actions truly lowered risk.
Preventing the Prediction System from Becoming a New Failure Source
Feedback loops abound: risk triggers scale-out, scale-out changes metric distribution, model sees risk drop, system scales back in, risk returns. Without cooldowns and action-state awareness, policies oscillate. Data feedback: model-triggered node replacement prevents the real hardware fault, so training data lacks positive samples; long-term training on only non-intervened data blinds the model to post-policy system behavior. Model, actions, and labels must share a unified event history, explicitly marking which outcomes were intervened.
Alert fatigue: risk prompts without owner, suggested action, and close conditions become ignored panels. Every prediction event type needs a service-level owner, response SLA, escalation path, and periodic retirement of rules/models with no proven benefit.
Safety controls fall into four boundaries:
Authority boundary: model has no direct production resource control.
Rate boundary: actions per time unit, per service, per failure domain are capped.
State boundary: pre-action checks current topology, capacity, release status.
Rollback boundary: every automated action has timeout, abort, and recovery conditions.
The prediction system itself needs drills: replay historical telemetry to verify old models reproduce past risks; inject feature latency, model unavailability, score spikes to confirm the system halts actions and falls back to base alerts. Drill goal is not to prove model “smartness” but to prove the whole chain fails safely.
Zero-to-One Implementation Roadmap
A pragmatic rollout in four phases, each delivering runnable capability:
Phase 1: Fault Catalog & Data Health Check
Pick 1–2 high-frequency, gradual faults with safe actions. Map entity identities, data sources, fault labels, action latencies. Replay recent months of events. Deliverables: prediction contracts, data quality report, baseline metrics—no model yet.
Phase 2: Rule-Based First Closed Loop
Encode expert experience into observable evidence, emit risk events, enter shadow mode. Complete deduplication, state machine, evidence display, result backfill. Even if rules are mediocre, this pipeline becomes the infrastructure for future models.
Phase 3: Temporal Validation of Model Gains
Compare rules, statistical baselines, and simple supervised models on the same event set. Focus on event-level recall, daily false positives, lead time, calibration—never random-split high scores. Only when model shows stable operational gain over rules does traffic expand.
Phase 4: From Human Confirmation to Limited Automation
Hand risk results to on-call, collect explanations and action feedback. Then gray-release reversible actions with global budgets and rollbacks. Every permission expansion requires a track record of benefit, false-positive rate, and safety.
“From zero to one” does not mean predicting all faults. It means the team has a trustworthy loop: they know what they predict, why, when to act, and how to fall back safely when the model breaks.
The first operable prediction loop is usually more valuable than the first high-score model.
Converting Lead Time into System Resilience
The most useful artifact of fault prediction is not “it will break” but a usable time window. Ten minutes may suffice to stop batch jobs; thirty minutes to migrate replicas; hours to swap hardware. Lead time only converts to lower impact when caught by an action chain.
At small scale, rules, statistical baselines, and human confirmation often suffice. As scale grows, incrementally add unified identity, streaming features, topology aggregation, model management, and action budgets. The evolution order must not be reversed: building a complex platform first, then hunting for problems to predict, easily yields an expensive system nobody trusts.
Next postmortem when someone asks “could we have caught it earlier?”, ask three sharper questions: Did stable evidence exist before the fault? How much lead time was needed to act? What is the cost of a false alarm? Answering these three moves fault prediction from idea into engineering.
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
Random Bulletin
17-year internet software developer specializing in AI applications, networking, architecture, and open source. Led the delivery of network services handling hundreds of millions of concurrent devices and tens of millions of QPS, and has three years of experience designing and building an agent platform. Follow to stay updated.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
