Beyond Tool Calls: FDE's Responsibility Chain for Production AI Agents
This article details FDE's framework for integrating AI agents into enterprise workflows, covering identity delegation, data egress inventories, action control with six record types, evidence binding across trace and audit layers, failure handling with four exception-path questions, compliance checkpoints embedded in state transitions, a five-party responsibility model, and ten required security deliverables.
The article opens by noting that a controlled runnable loop (from the previous installment) only reads data and generates evidenced risk tasks for human confirmation. The next step — letting an agent query inventory, contracts, and fulfillment data, then write back results after approval — introduces a full safety-responsibility stack: data egress, identity delegation, action confirmation, evidence retention, and failure handover.
Identity Inheritance: Agent Permissions Come from Delegation, Not Thin Air
Giving an agent a single service account is the easiest but most dangerous pattern; it merges different users, tasks, and business objects into one identity, making it impossible to answer who delegated, why it was allowed, or whether the authorization covers the current action. The author splits identity into three layers:
Initiating principal — every task must trace to a definite human user, service account, or automation chain. The record must link tenant, principal, role, initiation channel, and authorization context.
Delegation chain — the agent uses organizational permissions delegated per task and capability scope. Each hop must state who allowed reading which objects, who allowed calling which tools, which actions may only produce candidates, and which require a designated role's approval. Every hop needs scope, expiry, and revocation conditions.
Pre-submission real-time verification — identity valid at task start may be invalid at execution time. The source business system must re-check current principal, object state, permissions, action contract version, and approval records before committing a transaction. A disabled account, changed order status, modified parameters, or upgraded contract can all invalidate the original grant.
In the synthetic manufacturing case, planner A initiates a risk task; the agent queries orders and inventory within A's delegated scope. It can generate adjustment proposals but gains no right to modify customer commitments. When a proposal escalates to "adjust delivery commitment," it must enter a new risk level and approval chain; if inventory or parameters change, the prior approval expires. The article cites OpenAI Agents SDK's human-in-the-loop mechanism as proof that SDKs can support pause, approve/reject, and resume, but emphasizes that the SDK does not answer who in the organization may approve, what objects the approval covers, when it expires, or whether the source system must re-verify — those belong to the project's responsibility model and action contracts.
Data Authorization: Draw the Domain Inventory Before Discussing Model Capabilities
Teams often end data discussions with "enterprise data is not used for model training by default." The author argues this only addresses one purpose and does not equal data not being processed, retained, or logged, nor does it mean the customer approved data entering the model, connectors, observability platforms, or external tools. OpenAI's enterprise privacy page states business data is not used for training by default, but also clarifies that API input/output retention, zero-data-retention (ZDR), and specific security controls depend on endpoint, feature, eligibility, and configuration. Platform capabilities can serve as controls but cannot replace the customer's data classification, egress approval, and legality judgment.
Therefore, FDE must drive creation of a data-flow and egress inventory before integration. "Egress" includes not only cross-border or leaving the customer's data center, but also data leaving the original business system and established control domain to enter model context, tools, traces, logs, approval UIs, or third-party connectors. A minimum inventory must answer six questions (presented as a table):
What is the data? — category, sensitivity, source system, business owner.
Why use it? — business purpose, necessity, authorization basis, prohibited uses.
Where will it go? — model, tools, traces, application logs, approval UIs, external connectors.
In what form? — raw text, fields, summaries, masked values, controlled references, or statistical results.
How long retained? — retention periods, residency regions, and access permissions per endpoint and evidence layer.
How to exit? — deletion paths, credential revocation, handling when vendor is unavailable or data is mis-sent.
In the anomaly-order case, order IDs, inventory status, and contract-term summaries may enter the task context; unrelated fields like customer contacts should be minimized or masked; sensitive raw text should not be duplicated across logs just for "traceability." The model platform may retain necessary trace_id, version, and controlled summaries, while the original contract stays in the source system, accessible via controlled reference. The inventory is not a compliance sign-off artifact but the engineering input for configuring connectors, logging, masking, retention, and deletion policies. When the inventory cannot be clarified, the default action must be to shrink data scope or stop egress, not to integrate first and patch later.
Action Control: Tool Success ≠ Business Completion
Read-only queries and actions with business consequences cannot share the same control intensity. Querying inventory may need only object scope and field permissions; modifying commitments, creating formal tasks, sending customer notifications, or changing order status must go through full action contracts, human confirmation, and source-system verification chains.
Every business-consequence tool call must associate at least six record types (listed as an ordered list):
Action request — target object, key parameters, action contract version.
Identity verification — initiating principal, delegation chain, permissions, real-time status check result.
Decision basis — evidence references, rule or policy version, risk level, approval reference.
Execution identifier — idempotency key, executing system, transaction ID, or workflow number.
Execution status — received, executing, completed, failed, rolled back, status unknown, or human takeover.
Final receipt — source business system result, object changes, business events, audit references.
HTTP 200, MCP tool success, or a model claiming "already processed" only prove one layer responded; they cannot substitute the source system's confirmation of final business state. Asynchronous tasks may be awaiting approval, executing, or in unknown state. If the agent lacks the final receipt, it must not mark the task successful.
Human confirmation is not just an "OK" button. The approver must see the definite object, parameters, evidence, risk level, and action contract version, and the authorization scope must actually cover the action. If object state, parameters, evidence, or contract version change, the original approval must expire; rejection, timeout, withdrawal, and no-response must have defined outcomes.
In the synthetic case, the agent submits a candidate request "adjust order ORD-001 commitment date from A to B" with inventory, contract, and delay evidence. If inventory changes after approval, the action gateway must block further execution; source-system re-verification failure routes the task back to human handling instead of forcing write with stale approval.
Evidence Binding: Traces and Audits Are Not the Same
Many projects claim "full-chain observability" because they have model traces, SDK logs, application logs, and business-system logs. The problem: these records answer different responsibilities. A table contrasts four evidence layers:
Model/platform trace — answers which model, tool, and stage were invoked, where latency or failure occurred; cannot replace customer authorization and final business status.
SDK logs — answer how the agent paused, resumed, retried, or produced interrupts; cannot replace enterprise formal audit and source-system transaction records.
Application logs — answer task routing, policy decisions, UI interactions; cannot replace data authority and business result confirmation.
Source-system audit — answers who changed which object under what conditions and the final result; cannot replace model internal run details and engineering performance analysis.
These layers can be correlated via unified trace_id, task ID, object references, and transaction numbers, but must not be merged into a "big log" that substitutes one for another. Platform traces suit engineering observability; customer audit records reconstruct authorization and business responsibility chains; the source system remains the authority for transactions and final state.
Model invocation records should at minimum link: tenant and initiating principal, authorization context, model and prompt template versions, skill and tool contract versions, policy and evaluation versions, controlled references for inputs/outputs, human confirmations, failure retries, and the record's own access permissions, retention period, and deletion path. The article explicitly states it does not require saving raw model chain-of-thought, nor default copying of all sensitive inputs/outputs. Auditable focus is structured run facts, versions, and evidence references — not dumping more sensitive content into logs.
Failure Handover: Exception Paths Must Precede Automation Expansion
Safe agent integration cannot design only the happy path. Once in real workflows, common issues are not "model completely unavailable" but steps in uncertain state: model generated proposal but evidence expired, tool request sent but receipt lost, approval completed but object state changed, external platform unavailable yet unable to determine if transaction committed.
Each critical action must pre-answer four questions (bulleted list):
What signal triggers pause, and can the system prevent subsequent actions from cascading?
Who receives the exception, and what objects, evidence, versions, and execution state can the takeover person see?
Which failures allow retry, and how to use idempotency keys to prevent duplicate execution?
When state is unknown, who reconciles, and which source-system record is the final authority?
If the system cannot distinguish "not executed," "execution failed," and "result unknown," automatic retry may create duplicate actions, while direct rollback may overwrite already-successful business results. The professional conclusion is not to let the agent decide, but to pause, reconcile, hand over to human, and preserve the full event chain. FDE must write these paths into runbooks and complete controlled drills, but cannot decide business continuity for the business owner, classify security events for the security team, or modify final transaction state for the source-system owner.
Compliance Checkpoints Embedded in State Transitions, Not Pre-Launch Paperwork
The article references NIST AI RMF Playbook (Govern, Map, Measure, Manage) as voluntary guidance, not a mandatory checklist. For FDE projects, a more executable approach is embedding compliance checkpoints into state transitions. A table maps five project state transitions to minimum evidence:
Scenario initiation — business purpose, AI necessity, data classification, applicable requirements, prohibited scenarios, risk owner.
Prototype enters controlled loop — data flows, least privilege, test data boundaries, model/tool inventory, human takeover points.
Apply for production admission — security and abuse testing, permission review, retention/deletion policies, monitoring/alerting, incident response, rollback drills.
Enter stable operations — version changes, anomaly and privilege-escalation monitoring, periodic audit, residual risk review, user feedback.
Handover or exit — credential revocation, data and trace disposal, open risks, audit materials, responsibility transfer.
Article 06 only places these checkpoints and artifacts into the execution chain; production admission verdict, metric thresholds, and long-term business outcome proof are left for Article 07.
Five-Party Responsibility Allocation: FDE Does Not Assume Anyone's Liability
"Shared responsibility" cannot be operationalized without a concrete model. The author proposes a five-party model: FDE, Customer Business Owner, Customer Security/Compliance, Platform/Model Team, Source System Owner. It is not an industry standard; organizations may merge roles, but responsibilities must not disappear. A detailed table assigns responsibilities across five domains (data flow & egress, model invocation & trace, tool & action permissions, human confirmation, incidents & exceptions, production admission). Key points:
FDE models, implements, and continuously updates data flows; designs correlation fields and verifies replay; implements least privilege and contract verification; builds pause/display/expiry/resume mechanisms; establishes detection/pause/rollback/escalation paths; aggregates evidence but does not self-approve.
Customer Business Owner confirms necessity and purpose; decides auto/approve/prohibit; designates authorized approvers and owns decisions; decides business continuity handling; accepts business results and residual risk.
Customer Security/Compliance judges classification, egress, retention boundaries; defines access, retention, audit requirements; reviews permissions and segregation of duties; confirms risk grading and authorization rules; leads security incident grading and investigation; approves conditions or blocks.
Platform/Model Team explains processing and security capabilities; provides logs, traces, management capabilities; provides isolation, policy, observability capabilities; provides usable execution mechanisms; proves platform control items.
Source System Owner confirms authority, access, masking methods; provides source data and business record references; re-verifies approval applicability at commit; owns recovery, reconciliation, final state.
FDE's responsibility is not meeting coordination; it must participate in key integration design and implementation, verify checklists match runtime facts, drive missing responsibilities to explicit owners, and block expansion when boundaries are unclosed. But FDE cannot accept business/operational risk for the business owner, make compliance judgments or approve data egress for security/compliance, turn SDK auto-approval into human authorization, or let platform logs replace source-system responsibility for final state.
At Least Ten Security-Responsibility Artifacts Before Delivery
If an agent chain leaves only interface docs, prompts, and demo videos, FDE's safety integration is incomplete. Minimum artifacts (numbered list):
Data flow and egress inventory.
Identity, credential, and least-privilege matrix.
Model, prompt template, skill, tool, and policy version inventory.
Trace and business audit correlation fields.
Tool risk grading and human confirmation policy.
Sensitive action contracts and source-system final verification points.
Data retention, deletion, and exit disposal plan.
Security evaluation, abuse testing, regression baseline.
Monitoring, alerting, incident response, rollback, human takeover runbook.
Production admission approvers, blocking items, residual risk record.
These need not be ten separate documents; they can live in architecture decisions, configs, policies, tests, audit systems, and runbooks. But each must have a locatable artifact, owner, and version — not exist only in chat logs or an engineer's memory.
Pause integration or shrink scope when any of the following occurs (bulleted list):
Initiating principal and delegation chain cannot be reconstructed.
Data purpose, egress scope, retention, or deletion path lacks approval basis.
High-risk actions lack designated approver, or approval persists after changes.
Tool success cannot link to source-system final receipt.
No reconciliation and human takeover path for unknown state.
Platform capabilities are directly equated to customer compliance or risk acceptance.
Tool Calling Is Only the Start of Production Integration
Integrating agents into enterprise processes is not about giving models more permissions, but making permissions delegatable, revocable, and real-time verifiable; not stuffing more data into context, but ensuring every egress has necessity, scope, and disposal path; not accepting tool "success," but requiring confirmation, status, receipt, and handover for actions; not collecting logs, but enabling engineering observability, customer audit, and source-system facts to correlate without substituting each other.
FDE's core deliverable at this stage is the agent production integration boundary map and safety responsibility checklist. It owns engineering artifact completeness but does not accept business risk for the customer, approve data egress, make compliance judgments, or assume source-system transaction responsibility. Article 06 safely connects the agent into controlled real processes, but this is not a production admission conclusion. The next article will address: what hard gates must quality, security, operations, and governance satisfy; how to prove business results and user adoption over longer observation windows; and when the team should go live, expand, contract, or stop.
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
Data Bricklaying Diary
Records practices, thoughts, and pitfalls on the data grunt-work journey, sharing content on data platforms, data analysis, data processing, data governance, knowledge graphs, and more. Less theory, more hands‑on, making complex data technologies simple.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
