Agent Error Recovery: Why Retries Alone Fail — Idempotency & Reconciliation
This article explains why automatic retries after agent timeouts can duplicate business actions, and details how to distinguish call failures from executed side effects, implement execution-side idempotency guarantees, reconcile using original action identifiers, and define safe recovery rules with human escalation when outcomes remain unknown.
Failure of Call vs Business Action
An external call passes through request sending, service acceptance, business processing, and result return. The client-side exception does not directly indicate which segment failed. For example, a connection timeout differs from a response timeout after the request was sent; the latter may occur after the approval system has already completed submission. Even error codes require interface contract interpretation: they may signal validation rejection, processing failure, or a gateway that did not receive a downstream response.
When an exception occurs, first examine existing records to confirm how far the process reached:
Explicit rejection with guarantee of no business consequences — the action did not execute; correct the premise or end the task, do not blindly retry.
Accepted, returned approval instance reference — submission succeeded, approval process may still be running; track the original instance, do not resubmit.
No receipt or unclear receipt — execution uncertain; enter a pending-verification state and limit further progress.
This assumes the instance reference proves acceptance; if only a gateway queue number is returned, the actual acceptance result must still be queried. Submission success does not equal approval passed, and approval rejection does not equal tool call failure — the agent must not be allowed to resubmit until it obtains a result that satisfies the business intent.
Idempotency Is Not Just Adding an ID to the Request
Idempotency here solves: when the same business intent is submitted repeatedly, duplicate business consequences must not be produced.
Design: when a user confirms a purchase application version, the system generates a stable identifier for that submission and persists the identifier, submission content, and confirmation basis before the first send. On timeout or process restart, any resend must reuse the original identifier and content so the approval system recognizes it as the same action.
However, generating an identifier locally does not guarantee idempotency. The executing service must be able to recognize duplicate requests and, per contract, return the existing result, a processing state, or reject conflicting content.
When designing the interface, clarify:
Scope of identifier uniqueness (tenant, action type, resource).
Whether the same identifier with different parameters is explicitly rejected.
Whether concurrent identical requests still produce only one business result.
How long deduplication records are retained and whether task recovery may exceed that period.
Whether consistency between recording the idempotent identifier and producing business consequences is maintained if a failure occurs between them.
The last point is critical. If the service creates the approval and then separately saves the deduplication record, a crash in between can still cause duplicate creation. The execution side must link deduplication judgment with business execution via transactions, unique constraints, or other reliable coordination mechanisms.
Agent-side action records assist correlation and recovery, but local deduplication alone cannot guarantee the external system will not repeat execution. Idempotency guarantees must cover the full agreed consequences: no duplicate approval instance does not guarantee downstream notifications are not duplicated.
Additionally, if the user modifies the amount and reconfirms, that is a new business intent; the old identifier must not be reused for "anti-duplication." If the original action result is still unknown, it must first be verified whether it took effect to avoid old and new applications entering approval simultaneously.
Reconciliation: Clarifying What Actually Happened for the Same Action
Reconciliation here refers not only to financial reconciliation but to using the original action identifier and submission basis to verify the actual result with the authoritative system.
If the approval system returned an instance reference, query that instance; if the receipt was lost, look up by the pre-agreed client action identifier or idempotency key. Do not simply search by applicant and amount for a "similar" record and have the model decide it is the original submission.
Reconciliation must also verify the object, submission content, and version to avoid linking to another application. Receipts, event notifications, and active queries can complement each other, but state updates for the same action should be rule-constrained — late old messages must not overwrite already confirmed new results.
The most error-prone case is "query returns no result." The query interface may read a stale replica, the approval instance may still be asynchronously creating, or the original request may still be in the queue. Therefore, a single query finding nothing is usually insufficient to prove resubmission is safe.
Only if the interface contract confirms the original request was not executed and will not continue, or the execution side provides a valid idempotent replay guarantee, is there a basis for resubmission. Otherwise, retain the unknown state rather than fabricate certainty from an empty query result.
Retry Must Have Preconditions and Endpoints
Retry itself is not the problem. The problem is allowing the model to see an exception and then decide on unlimited retries, parameter-change retries, or switching tools to keep trying.
After an approval submission error, handling rules can be codified: before any possible re-write of business data, check permissions, input version, and whether user confirmation remains valid. Even with idempotency protection, expired authorization must not be reused.
The following table maps verification results to allowed recovery actions:
Confirmed accepted — associate original approval instance, enter result tracking.
Confirmed partial consequences but subsequent processing failed — retain completed parts record, continue per contract or compensate, do not re-execute entire chain.
Explicit rejection with no side effects — correct per rejection reason or end; business content change requires re-confirmation.
Result unknown but same request within valid idempotency protection — per interface contract, allow limited query or original request replay.
Result unknown but reliable query entry exists — continue query within verification period, pause dependent steps.
Confirmed original request produced no consequences, will not execute, fault retryable — limited retry while preconditions hold.
Verification expired still unknown, or lack safe recovery basis — block affected steps, hand to responsible person.
Example: if approval ended but result notification failed, the recovery action is the notification, not resubmitting approval. Unverified parts continue verification; confirmed consequences are retained.
These limits must be enforced by the execution program. After failure, gradually increase wait intervals with random offset to avoid thundering herds. Check SDK, tool layer, and Agent Loop — each layer may appear to retry only a few times, but stacked they can produce dozens of calls.
Within these constraints, the model can interpret exceptions, organize verification results, or suggest next steps; but whether execution is allowed is decided by the action contract and current evidence.
Exhausting retry attempts means stopping automatic attempts, not that the business definitely failed. Originally unverified results remain unverified.
Stopping the Task Does Not Automatically Withdraw Sent Requests
User cancellation or Agent process exit does not prove the approval system stopped processing. Requests already in the external queue may still create approval instances.
Therefore, cancellation should first prevent new business submissions, and a recovery process with appropriate permissions continues verifying sent actions. If withdrawal is needed, the approval system's explicit withdrawal capability must be used.
Compensation is similar: a new action for consequences that have already occurred, e.g., withdrawing a withdrawable application. It requires its own permissions, preconditions, action records, and result verification. An already approved application or one that triggered downstream purchase may not be withdrawable in the same way.
While the original submission is unverified, blindly initiating withdrawal may fail: the withdrawal request arrives first, the system finds no instance; later the original submission completes. Thus, "just compensate" cannot skip result verification.
When escalating to a human, the system should provide the original submission content, confirmation basis, queries performed, current unknowns, and allowed handling entry points. A human decision to continue execution does not mean rewriting unconfirmed history as "never happened."
Prove No Duplication Before Expanding Automatic Recovery Scope
This mechanism does not need to cover all tools initially. Start with one business action that has a clear query entry and reliable deduplication, and verify its behavior under fault conditions.
Purchase approval testing should at least cover:
External creation success, network break before receipt return — recovery associates original instance, external system produces only one submission consequence.
Concurrent requests with same identifier — execution side does not duplicate creation; different parameters explicitly rejected.
Query temporarily returns no result, original request later completes — system stays pending verification, does not create second instance based on that.
Idempotent record expired at recovery — do not treat old identifier as still having deduplication guarantee; escalate to verification or human.
Task cancelled, late success receipt arrives — retain real result, decide per business rules whether to withdraw.
Evidence comes from external business records, call traces, and control plane state changes — not just a final reply saying "recovery successful." Test environment passing only proves behavior under those fault conditions; production retention periods, permission configurations, and downstream interface semantics still need verification.
For simple read-only queries, focus on timeout, rate limiting, and retry budget; the full action coordination mechanism is not required. Conversely, if high-risk write operations lack reliable idempotency and a verifiable entry, limit the Agent to material preparation and preview generation, with a controlled business entry completing submission.
Summary
The value of recovery capability lies not in the Agent running more rounds after an error, but in its ability to clarify which consequences have already occurred and to stop affected steps when evidence is insufficient.
Actual troubleshooting follows three questions: Is this the original action? Would resending duplicate it? What does the business system record now? After clarification, decide whether to continue or stop.
When multiple Agents participate in a task, these boundaries must extend to subtasks. The next article discusses how to make them collaborate within explicit permissions and state boundaries, rather than relying on mutual dialogue retelling to advance the business.
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
Data Bricklaying Diary
Records practices, thoughts, and pitfalls on the data grunt-work journey, sharing content on data platforms, data analysis, data processing, data governance, knowledge graphs, and more. Less theory, more hands‑on, making complex data technologies simple.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
