Beyond Retry: Designing Fault Tolerance from Failure Models to Disaster Recovery

This article explains why retry alone is insufficient for fault tolerance, detailing how to classify failures via fault models, apply appropriate mechanisms like timeouts and circuit breakers, design recovery paths targeting state convergence, define RTO/RPO for disaster recovery, and validate all assumptions through fault injection drills.

Data Bricklaying Diary
Data Bricklaying Diary
Data Bricklaying Diary
Beyond Retry: Designing Fault Tolerance from Failure Models to Disaster Recovery

The previous article discussed observability: making business results, state changes, dependencies, and resources continuously visible. Once a failure is observed, the next question is how the system limits impact and recovers to a correct state.

Many systems' first reaction is retry. Connection failed? Retry. Request timed out? Retry. Message processing exception? Retry again. Retry can handle some transient failures, but it cannot determine whether the previous operation already took effect, nor does it automatically fix business actions executed multiple times.

Fault tolerance is not "try again when error occurs"; it is first understanding what failures can happen and what they affect, then designing detection, isolation, recovery, and manual intervention paths.

Timeout Only Means No Result Received

Continuing with the refund example from earlier.

A refund task submits a request to the payment channel; after five seconds no response arrives. To the local program this is a timeout; to the business there are at least three possibilities:

The request never reached the channel; refund not executed.

The channel already completed the refund; the response was lost on the way back.

The channel is still processing; result not yet determined.

The client cannot distinguish these three cases from a single timeout. If it immediately resubmits and the channel lacks stable idempotency, the same refund may execute twice.

Call timeout != Business failure
Return success != Business state already finally correct

Therefore, without additional evidence, timeout must be treated as result unknown. The system should retain the business operation ID, idempotency key, channel request number, and current task state, then confirm the result via channel query or reconciliation before deciding whether to resubmit.

Establish Failure Model First, Then Choose Handling Mechanism

A failure model is not a list of all exception class names; it describes where the failure occurs, whether the business result is determined, how impact propagates, and which party can safely recover.

In the refund chain, at least four failure types must be distinguished:

Transient Failure : Typical Situation - Brief network jitter, temporary connection refusal, short-term overload. Handling Principle - Bounded retry when failure is likely recoverable and operation can be safely repeated.

Permanent Failure : Typical Situation - Parameter error, insufficient permission, business condition not met; will not self-recover under current conditions. Handling Principle - Stop immediately; correct request, permission, or business decision.

Partial Success : Typical Situation - Channel already refunded, but local state or subsequent event update failed. Handling Principle - Preserve checkpoint; recover unfinished steps; do not repeat already effective actions.

Result Unknown : Typical Situation - Request timeout or receipt lost; cannot confirm whether channel executed. Handling Principle - Stop blind resubmission; query, reconcile, or escalate to manual.

The same technical exception may belong to different types in different businesses. Cache read timeout may degrade to database read; fund operation timeout may mean result unknown. Failure classification must combine business semantics, not decide automatically based only on HTTP status codes or exception types.

A complete failure model also records: failure point, detection signal, propagation scope, user impact, safe action, recovery target, and responsible party. These should enter architecture decisions, state models, and runbooks — not be decided on the spot when a production failure occurs.

Retry Only Suits Some Transient Failures

Safe retry requires at least four conditions: failure is likely short-lived; repeating the operation causes no extra side effects; enough time remains for another attempt; and retry does not make an already overloaded dependency harder to recover.

Microsoft's retry pattern documentation also emphasizes that retry targets transient faults and must consider whether the operation is idempotent; for long-lasting faults and non-transient exceptions, repeated attempts only waste resources.

A controlled retry strategy must define:

Which errors are retryable and which must stop immediately.

Maximum attempts, total time budget, and per-attempt timeout.

Progressively increasing intervals with small random jitter to avoid many instances retrying simultaneously.

Which layer is responsible for retry, avoiding client, gateway, and service layers stacking retries.

What happens after retries are exhausted: fail, query, degrade, or manual handling.

Downstream single timeout and retry must be constrained by upstream remaining time budget; each layer cannot independently consume a full timeout budget. The caller giving up waiting does not mean the remote operation has stopped; without deadline propagation and cancellation semantics, the background may continue executing after the user receives failure.

Assume client, business service, and channel adapter each "execute at most 3 times including first attempt"; the bottom layer could be invoked 27 times for one business request. Layered retries do not add threefold fault tolerance — they amplify failure into a retry storm.

Timeout, Circuit Breaker, and Isolation Only Limit Impact

Fault tolerance mechanisms have different responsibilities; they cannot be merged into a single "enable for reliability" configuration.

Timeout : Primary Problem Solved - Limit time and resources occupied by a single wait. Cannot Prove - Remote operation definitely did not execute.

Retry : Primary Problem Solved - Give some transient failures a chance to recover. Cannot Prove - Repeated execution definitely has no side effects.

Circuit Breaker : Primary Problem Solved - Quickly stop calls when dependency continuously fails. Cannot Prove - Dependency or historical tasks have recovered.

Isolation : Primary Problem Solved - Prevent a class of tasks, dependencies, or tenants from exhausting shared resources. Cannot Prove - Isolated business can stop being processed.

Rate Limiting & Backpressure : Primary Problem Solved - Prevent new work from crushing system and dependencies. Cannot Prove - Already accepted work has completed.

Degradation : Primary Problem Solved - Preserve core business when some capabilities are unavailable. Cannot Prove - Degraded result is equivalent to normal result.

These mechanisms mainly protect the system during failure; they are not responsible for automatically returning business state to the correct result. A circuit breaker resuming calls does not mean the backlogged refunds during the break have been processed.

Degradation must also respect business semantics. Temporarily returning empty results for product recommendation may be acceptable, but when the refund channel is unavailable, a fake "processed successfully" cannot serve as a degraded result.

Recovery Target Is State Correctness, Not Just Service Restart

After a failure ends, restarting instances and cutting traffic back to the primary site only means the service may accept requests again. In-flight tasks during the failure, unknown results, duplicate deliveries, and state discrepancies still need handling.

For the refund chain, a complete recovery path might be:

Restore service access
-> Inventory in-flight tasks and unknown results
-> Query channel authoritative state by business identifier
-> Recover unfinished steps or execute compensation
-> Reconcile and correct discrepancies
-> Confirm backlog, discrepancies, and user impact have converged

Compensation is not distributed rollback. Already effective external actions usually cannot be undone like database transactions; they can only be corrected by new business actions. Compensation may also fail, so it must have its own idempotency key, state, retry boundary, and manual entry.

Recovery completion judgment must start from user and business results: can the service accept new requests? Is historical backlog cleared? Are unknown results confirmed? Is data consistent again? Are affected users handled?

Disaster Recovery Must First Define RTO and RPO

Previous sections focused on component and dependency failures. When failure scope expands to data center, availability zone, site, or entire region, disaster recovery design is needed.

RTO (Recovery Time Objective) is the maximum acceptable time from business interruption to recovery of critical capabilities. RPO (Recovery Point Objective) is the maximum data loss the business can tolerate. The AWS Well-Architected Framework also uses these two objectives to distinguish service recovery time and acceptable data loss.

RTO: How long after failure must critical business capability be restored
RPO: How much data can be lost at the moment of failure

Different businesses need different targets. Refund acceptance, refund execution, historical query, and analytical reporting have different interruption impacts; they should not all be set to the highest level. Shorter RTO and RPO closer to zero usually require higher redundancy, replication, and operational cost.

Backup does not equal disaster recovery. Backup only proves a data copy existed at some point; it cannot prove that data can be restored within the target time, nor that applications, configurations, secrets, messages, external dependencies, and traffic switching are ready. Synchronous replication may also replicate accidental deletions, erroneous data, and malicious operations to the DR site.

Disaster Recovery Switch Can Also Amplify Failure

Failover is not simply pointing traffic to another address. The standby site must have usable business data, configurations, identities and secrets, message offsets, dependency connections, and capacity.

Before switching, at least these questions must be answered:

Has the primary site stopped writing, or could two simultaneous authoritative writers form?

How far behind is standby data; does it meet RPO?

How are in-flight tasks and unacknowledged messages handled?

Can external channel callbacks, gateway routing, and user connections redirect to standby?

How to verify business correctness after switch, and under what conditions to switch back?

If primary site state is unclear, standby data integrity unconfirmed, or both sides might write externally simultaneously, automatic failover may turn a partial outage into data divergence. At that point, automatic progression should stop; an authorized decision-maker must decide based on recovery targets and current evidence whether to continue switching, maintain degradation, or enter manual recovery.

What Architecture, Platform, and Service Teams Are Responsible For

Fault tolerance and disaster recovery cannot be left only to ops teams, nor stay only in architecture documents.

Architecture Design : Main Responsibility - Define failure model, propagation boundaries, recovery targets, state convergence methods, and manual intervention conditions.

Platform Capability : Main Responsibility - Provide timeout and retry baselines, traffic isolation, health checks, backup and restore, fault injection and drill tools.

Service Team : Main Responsibility - Maintain idempotency, checkpoints, compensation, reconciliation, runbooks, and on-call response; responsible for business state recovery.

Business owners must also confirm: which capabilities must recover first, which can degrade or delay, how much interruption and data loss are acceptable, and who has authority to decide switch, compensation, and recovery.

When failure type cannot be confirmed, external result unknown, automatic recovery repeatedly fails, compensation would create new high-risk side effects, or standby data cannot meet RPO, the system must preserve the scene and escalate to human decision — not cover the problem with another round of retries.

Fault Tolerance Conclusions Must Pass Fault Drills

Code review and config checks can find some issues, but cannot prove the system can recover under real failures.

Transient failure can recover safely : Verification Goal - Retry completes business, no duplicate side effects or timeout amplification. Evidence Source - Fault injection, idempotency verification, call records.

Dependency failure does not drag down whole system : Verification Goal - When dependency continuously fails, resources isolated, core business still available. Evidence Source - Circuit breaker, isolation, and overload drills.

Unknown result can converge : Verification Goal - Whether channel executed or not, both return to single correct state. Evidence Source - Drop receipt, channel query, reconciliation result.

Node loss does not lose tasks : Verification Goal - Executor exits at any checkpoint; task recovers and does not repeat completed actions. Evidence Source - Process termination, task recovery, state validation.

Backup and DR are usable : Verification Goal - Data and service restore within RTO/RPO; business correct after switch. Evidence Source - Backup restore test, DR switch drill, business reconciliation.

Team can complete handling : Verification Goal - Non-original author can identify failure, mitigate impact, and recover business. Evidence Source - On-call drills, runbooks, incident records.

Fault drills are not random destruction of production systems. They need explicit hypotheses, impact scope, abort conditions, observation metrics, and recovery steps; verify first in controlled environments, then expand scope based on business risk.

After drills, record failure detection time, impact scope, mitigation and recovery time, actual data loss, human decisions, and unfinished remediation. Retrospective purpose is not to assign all blame to one person, but to feed newly discovered failure combinations and recovery gaps back into architecture, platform, and runbooks.

Agent Tool Timeouts Also Cannot Be Blindly Retried

Agent calling a clearly read-only query tool with no external side effects that times out can usually be retried under bounded conditions; calling a tool that modifies funds, permissions, or business state that times out may have already produced external side effects.

At that point, the Agent should not automatically re-execute based on a single tool error; instead it should retain task ID, tool call ID, idempotency key, and external receipt, and query the authoritative source for true state. When confirmation is impossible, the task should enter blocked or wait for human decision, not let the Agent guess a result with more calls.

Small Systems Also Need Minimum Recovery Path

Internal systems with few dependencies, lower business risk, and recoverable manually during working hours do not need multi-site DR, complex circuit breaker platforms, or large-scale fault drill systems from the start.

Minimum fault tolerance design still requires recording: which operations can retry, which failures must stop immediately, how data is backed up and restored, who handles failures, and how business confirms it has returned to correct state.

The timing for introducing complex mechanisms should be decided by real failure models, recovery targets, and drill results — not because some architecture template lists those components.

Summary

Fault tolerance is not adding retry at every layer; it is limiting impact scope and letting business state eventually return to correct result based on failure type and business semantics.

First judge what failure occurred,
Then decide retry, isolate, degrade, compensate, or escalate to manual.

Timeout cannot prove remote operation did not execute; service restart cannot prove business has recovered; backup existence cannot prove system can rebuild within RTO/RPO. These conclusions must pass fault injection, state validation, backup restore, and DR switch drills.

Next article will continue discussing security architecture: how identity, secrets, sensitive data, and audit are designed from system boundaries.

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

distributed systemsdisaster recoveryfault toleranceRetryidempotencytimeoutCircuit Breakerfault injectionisolationRPORTOdegradationstate convergence
Data Bricklaying Diary
Written by

Data Bricklaying Diary

Records practices, thoughts, and pitfalls on the data grunt-work journey, sharing content on data platforms, data analysis, data processing, data governance, knowledge graphs, and more. Less theory, more hands‑on, making complex data technologies simple.

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.