Designing a Resilient AI Runtime: Observability, Degradation, Takeover, Rollback
This article outlines a production-grade AI runtime architecture that separates processing state from business conclusions, enforces bounded model execution with retry budgets and circuit breakers, ensures idempotent external actions, defines explicit degradation paths, structures human-in-the-loop workflows, extends observability to judgment quality, and validates rollback and failure injection before launch.
1. Separate Processing State from Business Conclusion
Many issues stem from an overly simple status field. If a task only has "success" and "failure", a video that is genuinely non-compliant gets mixed with a video that never finished analysis. An HTTP 200 may only mean the request was accepted, not that the model produced a parsable answer.
The author proposes splitting state into two dimensions:
Processing state – where the workflow has reached: accepted, processing, completed, failed. "Completed" means this processing run ended, not that the business loop is closed.
Business conclusion – what judgment the evidence supports: pass, fail, pending-review, unknown. "Unknown" means insufficient evidence; "pending-review" means the system has escalated to a human.
This is an example contract; other domains (e.g., security) can use conclusions like "suspected", "confirmed", "dismissed".
With this split, code can enforce constraints: when processing fails, the system must not emit an automatic "pass" conclusion.
Two states are still not enough. The system must also record the next action: wait-retry, request-more-evidence, escalate-to-human, query-external-result, or terminate. State describes "what happened"; action describes "what to do next". Neither should be left to the model's improvisation.
2. Put the Model Inside a Bounded Runtime Architecture
A controllable system cannot be just "input → LLM → output". The full main chain is: task acceptance, input validation, inference execution, result verification, business decision, result delivery. Alongside it sit a task state store, evidence log, exception handling, and a human workbench.
Three responsibilities must be separated:
Model execution layer – responsible for "understanding". It extracts facts, recognizes scenes, proposes candidate conclusions, but its output is only material for verification.
Policy control layer – responsible for "deciding". It validates fields, rules, evidence locations, permissions, and decides whether to continue, stop, review, or allow a business action.
Delivery execution layer – responsible for "making it effective". It calls downstream APIs, records delivery results, handles duplicate requests, and distinguishes sent, received, and processed. Generating text "alert sent" does not equal a successful delivery.
The same applies to autonomous agents: they may propose new investigation steps, but every tool call must pass permission, parameter, budget, and approval checks. The executor should verify task version and expiry to avoid cancelled tasks triggering stale actions.
This architecture makes the process more controllable but does not guarantee the model's semantic judgments are always correct. Model quality still requires evaluation, sampling review, and continuous improvement.
3. Not Every Failure Deserves a Retry
Transient faults: retry with a budget
Network jitter, brief rate-limiting, and temporary service errors can be retried when the interface contract allows. Retries should use backoff, jitter, and respect the task's total time budget and total attempt count. They must also account for retries already performed by lower-level SDKs to avoid multiplicative explosion. Microsoft's retry pattern documentation also stresses distinguishing transient from non-recoverable faults.
The key is not how long a single call can wait, but how long the whole business can wait. Real-time alerts and offline video reports obviously cannot share the same time budget.
Input errors: raw retry is pointless
Corrupted files, missing required fields, unsupported formats should enter a repair or evidence-supplementation flow. Sending the same undecodable video to the model ten times will not make the file whole.
Semantic uncertainty: don't retry until you get the answer you want
Dark footage, occluded actions, conflicting evidence indicate insufficient grounds for judgment, not necessarily an interface fault.
If the model first answers "cannot determine", repeatedly rephrasing the prompt until it says "pass" is not resilience – it is cherry-picking the desired answer.
Adding evidence, limited independent re-reviews, or using a verified fallback analysis path can improve judgment. Multi-model agreement cannot substitute for evidence itself.
Persistent faults: stop hammering a broken dependency
When a downstream service fails continuously, the system can pause calls at a preset threshold, switch to a degraded path, and use a few probe requests to confirm recovery. This is the purpose of a circuit breaker; it does not replace business exception handling, nor does every chain need its own implementation. Microsoft's circuit breaker pattern distinguishes the roles of circuit breaking and retry.
Resource isolation is also needed: batch video analysis must not exhaust the concurrency quota reserved for real-time alerts, and the human review queue must not become an unbounded "black hole".
4. The Most Underestimated Failure: The Action May Have Succeeded, You Just Didn't Get the Result
Suppose the system sends a notification to an alert platform. The platform receives it, but the response is lost due to a network break. The caller sees a timeout but cannot conclude "not sent". If it immediately retries with a new request ID, duplicate alerts may be generated.
Timeout describes the caller not receiving a result in time; it does not necessarily mean the external action did not happen. For side-effecting interfaces, an idempotency contract or a result reconciliation mechanism is required. AWS's idempotent API documentation discusses this retry risk.
Engineering design can handle this as follows:
Before sending, record the action intent and assign a stable action ID; retries of the same action reuse this ID.
The receiver handles duplicate requests per agreed scope, parameter consistency, and retention period.
On timeout, mark delivery state as "pending-reconciliation", query the receiver's records or wait for a trusted receipt.
Retry under idempotency guarantees; if confirmation is impossible and safe retry conditions are absent, escalate.
If a local transactional outbox table is used, the business state and the pending-delivery record should be committed in the same transaction, then a sender delivers. This only solves persistence within a specific boundary, not end-to-end "exactly once". Downstream duplicate processing must still be constrained.
Two kinds of IDs must be distinguished: a technical retry's action ID should stay stable; a new event from the same camera later must not be permanently deduplicated as the old event. Genuine new actions, content revisions, and risk escalations all need clear version or event semantics.
5. Degradation Is Not Just Swapping Models – It's Pre-Agreeing on What to Do Less
"Primary model unavailable, automatically switch to backup" sounds natural. But different models may understand the same video differently, produce different structured outputs, locate evidence differently, and apply different thresholds. Unverified automatic switching can turn an explicit service fault into a hidden judgment error.
A degradation plan must first answer: what guarantees does this path still provide? What does it not guarantee? Which results must be marked as degradation-produced? When do we exit degradation?
Engineering video inspection: delay the report, never fake completion
Assume the system checks whether critical construction steps were completed per requirements.
When the model is temporarily unavailable, deterministic checks (file integrity, duration, resolution) can still run; checks requiring visual understanding stay "unknown" and wait for recovery or human review.
If a key step lacks usable footage, the system should show "this item cannot be judged", not fill the gap with passes from other items. If business rules mandate blocking downstream flow, an explicit rule should execute while preserving "insufficient evidence" as the reason – not masquerading as a discovered violation.
Partial results must be explicitly tagged: which items completed, which are missing, whether conditions for a final report are met.
Intelligent security: tolerate missing explanations, never silently erase risk signals
Assume the existing detection chain identifies candidate intrusion events, and the LLM further explains the scene to reduce false positives.
When the LLM is unavailable, whether to continue emitting candidate alerts should be decided by a pre-reviewed risk policy. The system must not turn a candidate event into "no anomaly" just because the explanation step timed out.
One optional design keeps the basic detection and notification channel, sending "pending-confirmation events" to on-duty staff with a clear note that enhanced analysis is temporarily unavailable, while retaining key clips. Deployments without independent detection capability should explicitly report impaired monitoring and activate appropriate duty measures.
This does not mean blindly pushing all candidates. Deduplication, aggregation, prioritization, and escalation strategies remain important, and must avoid merging new high-risk events into existing ones.
The above scenarios are architectural examples, not universal site-safety procedures. Systems involving personnel safety or equipment actuation require risk assessment and dedicated validation by responsible parties; an LLM cannot be the sole protective layer.
The bottom line of degradation is not "the page still loads", but that business-critical constraints still hold and lost capabilities are explicitly communicated.
6. "Escalate to Human" Is Not the Endpoint – It's Another Process That Must Be Designed
Putting a task into a list does not guarantee anyone will handle it.
A truly usable human takeover process must define four things:
Who handles it – dispatch by specialty, region, event type, or on-call schedule. When no one claims it or it times out, there must be an escalation target, not just a growing backlog.
What they see – the workbench should surface the raw clip, timestamp, triggering rule, model candidate conclusion, uncertainty reason, and handling history. Reviewers should not have to hunt through tens of minutes of video for a vague question.
What they can do – confirm, overturn, request more info, or decide to terminate; different actions map to different permissions. High-impact actions should have pre-defined secondary approval.
How to avoid conflicts – once claimed, a task has clear ownership and a timeout reassignment mechanism; submission validates state version. A human-confirmed conclusion must not be overwritten by a late model response. Final state transitions should be done by server-side conditional updates or transaction checks, not just by greying out a button.
Original model results, human overrides, and final effective results must all be stored separately. Preserving differences enables accountability and reveals issues in rules, models, and input materials.
Human corrections can become improvement signals, but should not enter training or rule updates without sampling, authorization, and version control. A one-off handling does not necessarily represent a universally valid new rule.
7. Observability: From "Server Healthy" to "Business Has an Owner"
No server alarms does not mean tasks aren't stuck.
Basic monitoring usually watches latency, traffic, errors, and saturation – the four golden signals from Google SRE . AI systems need to extend observation points into the judgment and delivery chains.
The author suggests three monitoring groups:
Processing chain: oldest task wait time, per-stage latency, timeout rate, retry amplification factor, dependency availability, per-task cost.
Judgment quality: structure validation failure rate, evidence missing rate, unknown ratio, review ratio, and human override rate with sample size and review scope.
Business closure: delivery success rate, pending-reconciliation action count, unclaimed task count, overdue-unhandled count, event confirmation and resolution time.
Average values alone are insufficient. A few tasks stuck for a long time can be hidden by a pretty average latency; fewer alerts may mean the chain is broken, not that the environment got safer.
Rising unknown ratio does not automatically prove model degradation. Input quality, scene distribution, rule and version changes can all be causes. Alerts provide investigation leads, not conclusions.
A single task's evidence chain should at least link: task ID, stage and attempt IDs, input version, model and prompt versions, rule version, tool calls, evidence references, final state, and external delivery receipt.
Tracing does not mean dumping all raw content into ordinary logs. Videos, screenshots, and detailed inputs must be managed by authorization, de-identification, and retention policies; logs keep necessary references and summaries with controlled access.
Finally, every alert that wakes a human must state: what business is affected, which tasks are covered, who is responsible, and where to look next. Otherwise, the flashier the dashboard, the easier the real problem is to miss.
8. Rollback Is Not Just Reverting the Model – It Handles Three Objects
After a new version rolls out, if pending-review volume spikes, clicking "rollback" only solves part of the problem.
Three objects must be handled:
Future new tasks – shift traffic back to a verified, compatible old version combination. This combination includes not just the model, but also prompts, rules, knowledge, preprocessing code, and output contracts.
In-flight tasks – decide whether to continue with the original version, cancel, or re-execute. Each attempt must fix the version combination; a single report must not use new rules for the first half and old rules for the second. Before re-execution, check whether external actions have already been produced.
Already-effective results – list affected tasks, reports, and alerts; assess whether re-review, correction, or compensation is needed. Rolling back the version does not automatically recall sent notifications or erase actions already taken by humans.
This explains why version records and action records must be in place before launch. Without records, it is hard to accurately scope impact after an incident.
For incompatible data structures or interface changes, do not assume the old program still runs. Rollback paths need prior validation; when direct rollback is impossible, use a reviewed fix or migration plan.
After the issue is resolved, turn the failure samples into fixed regression test cases. The post-mortem artifact should not be just "be more careful next time", but new validations, monitors, permission constraints, or failure drills.
9. Before Launch, Let the System Break on Purpose
A complete architecture diagram does not mean the failure paths actually work.
In isolated or explicitly authorized, scope-controlled environments, run failure drills:
Make the model continuously timeout: does the total time budget enforce? Are resources released? Does the task enter the agreed state?
Return unparsable answers: can the system detect a contract violation instead of stitching together a "normal-looking" result?
Lose the delivery receipt: does it enter the reconciliation flow? Does it avoid duplicate external actions?
Leave the human queue unclaimed: can it escalate by priority and SLA instead of waiting forever?
Deliver an old model result after human submission: is the final conclusion still bound by version and state constraints?
Switch versions while some tasks have already taken effect: can the system distinguish new tasks, in-flight tasks, and affected results?
Drill acceptance criteria must be written in business language: no false passes, no silent losses, no unauthorized actions, every task needing handling can find an owner.
Then use logs, state records, and delivery receipts to prove these conditions actually hold.
Conclusion: A Mature AI System Has Clear Failure Modes
Moving from demo to production, the biggest change is not that the model got a little smarter, but that the system starts bearing real consequences.
Therefore, when designing AI systems, ask less "what else can it automate?" and ask more:
When it doesn't know, will it honestly stop?
When an external action's result is unclear, will it reconcile first?
When a critical capability fails, is there an explicit degradation plan?
After human intervention, will the system respect decisions already made effective?
The model is responsible for making the best judgment it can; engineering is responsible for ensuring that the uncertainty of that judgment is not silently amplified into business risk.
Observability, degradability, takeover-ability, rollback-ability are not add-ons to bolt on after launch.
They are part of the AI business system itself.
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
