Output Isn't Completion: The Verifiable Evidence Gap in AI Agents
This article argues that AI agents often mistake model output or tool calls for task completion, but real business completion requires verifiable read-back evidence, idempotency handling for timeouts, explicit acceptance criteria, and structured completion records — not just a 'final_output' signal.
Many AI Agent architecture diagrams end with a box labeled Output Result . The user sets a goal, the Planner breaks it down, the model reasons, Tools and MCP connect to external systems, Memory retains context, and Reflection checks the work. Arrows flow to a final answer, a file, or an external operation — and the task appears done.
In a real business system, that box is not enough. Consider a task:
Summarize recent product changes from three competitors into a brief and post it to the project Slack channel.
The Agent searches, compiles, generates the brief, and calls the Slack API. The request times out. At this point you cannot judge success or failure — the message may not have been sent, or it may have been sent but the response didn't return. If the Agent replies "task completed", the user sees nothing in the channel; if it treats the timeout as failure and retries immediately, the channel may get two briefs. The model produced output, the tool was invoked, but the system has no evidence that the task is actually complete.
The missing judgment: an Agent must not only produce content, it must know which facts are sufficient to prove it can stop.
Completion Has Multiple Layers
Previous articles in this series decomposed Planner, Tool, MCP, Memory, Harness, and Handoff — handling goal understanding, action selection, system connection, and next-round scheduling. Wiring these together still does not equal "task completed." In engineering, completion has at least these layers:
Model has output — proves this generation round ended; does not prove user goal achieved.
Tool call finished — proves external capability returned a result; does not prove returned result is true and usable.
Artifact generated — proves file, report, or record exists; does not prove content meets requirements or others can access it.
External action took effect — proves message, order, or config can be read back; does not prove business object and result are fully correct.
Goal passed acceptance — proves success conditions have matching evidence; only now is it appropriate to declare completion.
The OpenAI Agents SDK run loop illustrates this boundary: when the model produces compliant text and stops making tool calls, the Runner treats it as final_output and the agent loop ends. This is a run termination condition , not a business completion condition .
Looking at Coding Agent source code, the same boundary appears: assistant, tool_result, result, status, and task_notification are distinct messages; Stop, TaskCompleted and similar events record the run lifecycle. The framework can log what a tool returned and how a round ended, but it cannot judge whether the deliverable was correctly received by the business. final_output only says this run has closed. It cannot answer: was the brief posted to the right channel? Does the link have permissions? Does the content cover the three specified competitors? The framework isn't wrong — it simply doesn't know what each business considers "delivered." Completion conditions must be supplied by the application.
Reflection Is Not Observation
In some Agent architectures, Reflection sits after action: act, observe, reflect, then decide whether to continue. Its effectiveness depends on what Reflection receives.
If a tool only returns: message sent successfully The model can at most judge "the tool says it sent successfully." It cannot see whether the message actually appeared in the target channel, nor whether the report link inside the message is accessible.
Asking the model to reflect again will not conjure missing facts; it will likely just rephrase the tool's claim.
Therefore we must look at observation . After an external action, does the system re-read the real state?
Send a message → obtain channel and message ID → read back by that ID.
Write a file → check existence, readability, required sections.
Modify config → confirm the running system has loaded the new value.
Create an order → verify order status, amount, and object.
These checks need not be done by the model. Deterministic code is better suited: query by message ID, compare content digests, validate fields, check state transitions. kimi-code 's tool documentation includes a rule that fits here: after a successful Edit or Write, you don't have to re-read the file just to prove the write happened; but if the task depends on a precise file, API, or output shape, you must verify the final external contract before finishing.
The same applies to Agent systems. Not every step needs verification, but the task endpoint must be verified. The closer to the endpoint, the more the judgment should rely on read-back evidence rather than the model's self-confirmation.
Define the Endpoint Up Front
If you only ask "is this complete?" at the end, the Agent can only cite its own existing output as proof.
The planning phase can directly write down completion conditions. It doesn't need to become a full workflow; it just needs to state three things: what result to obtain, where the result lands, and what facts prove it.
Back to the Slack brief example, completion conditions can be stated plainly:
The brief is not "generate some text" but a document covering three competitors, a specified time range, and sources.
The final object is not a temporary file in the Agent's working directory, but a message in the project Slack channel.
Proof of completion: a stable document link, target members can access it, Slack returns a message ID, and reading back by that ID confirms channel, body, and link all match.
The Planner's task breakdown then shifts from "search, write, send" to a closable chain:
Goal → Action → Observation → Verification → Complete
└→ Continue / Wait / Escalate to humanVerification need not pursue absolute perfection. For a weather query, a valid tool result suffices; for a draft, check file existence and structure; for payments, permissions, publishing, or deletion, evidence must be concrete — readable, comparable, and sometimes human-confirmed. The higher the failure cost, the less "completion" can rely on a single success response.
Timeout First Means Unknown
Some systems give the Agent only two choices: success or failure. With many external actions, that binary is insufficient.
Back to the Slack message: a timeout only means the caller didn't receive a response in time; it does not prove the action didn't happen. The system should enter a third state: Unknown .
Unknown is not failure. It means the system lacks evidence to declare completion, and must not blindly retry.
Idempotency design solves exactly this. AWS's "Making retries safe with idempotent APIs" describes a classic scenario: a create-resource request hits a network timeout; the caller cannot know if the resource was created; a naive retry may create duplicates. AWS's solution: the caller provides a unique request identifier. The server uses this identifier to recognize the same intent and returns semantically consistent results on retry. Stripe's idempotency keys follow the same principle.
In an Agent system, this request identifier can come from task_id, step ID, and action sequence number. The first call and any retry use the same idempotency key ; after a timeout, query by that key — if an effected result is found, proceed to verification; if confirmed absent, retry.
If the external tool supports neither idempotency keys nor a reliable query interface, the Agent should not pretend to know the answer. Preserving unknown, handing existing evidence and risks to a human, is more credible than feigning certainty.
A reliable Agent not only retries — it also knows which actions must not be repeated casually.
Acceptance Is Not a New Agent
Some architecture diagrams add a Verifier Agent. But adding a role does not mean the completion judgment is solved.
I'm cautious about this approach. If acceptance is just checking file existence, querying a message ID, or comparing database fields, inserting another model only adds a layer of uncertainty. When completion conditions are clear, code, rules, and state machines are more direct and easier to test.
Only when the result itself requires semantic judgment does model-based acceptance add value. For example: checking whether a competitor brief misses important product changes, whether different sources contradict each other, or whether conclusions exceed the evidence.
In that case the model can participate in evaluation, but the materials it references, the scoring criteria, and the final evidence must be recorded.
What we need to add is acceptance capability , not necessarily an acceptance Agent. It can be a closed loop inside the Runtime.
Understand goal, handle ambiguous requirements → Model
Execute actions, record request identity → Runtime / Tool
Read back external state, validate exact fields → Deterministic program
Evaluate content quality and semantic differences → Model or human
Persist evidence, decide stop or continue → Runtime
This division avoids a common pitfall: letting the model execute, self-evaluate, and then declare itself passed.
The model might be right. But when something goes wrong, it's hard to tell whether the error was in the action, the observation, or that premature "completed" statement.
Keep the Record Small
A usable completion record doesn't need to be complex. At minimum it captures task status, success criteria, external result identifiers, verification result, and the next action.
{
"status": "completed | failed | unknown | needs_human",
"success_criteria": ["report_created", "message_visible"],
"result_ref": "slack:channel/message_id",
"idempotency_key": "task_id/step_id/action_index",
"verification": "matched",
"evidence": ["report_url", "message_id"],
"next_action": null
}These fields look less glamorous than Planner, Reasoning, or Memory, yet they determine whether the Agent's output can enter the next business process.
Downstream systems no longer need to parse a large block of natural language to guess if the task ended. Monitoring can distinguish completed, definitively failed, and result-unknown. When a human takes over, they see already-verified facts, not a vague error message.
If I redraw that Agent architecture diagram, I'd leave a little more space under "Output Result."
Output is merely the model finishing its expression; a tool return is merely one call ending. Only when the result is observed, verified, and mapped back to the user's original goal does the task earn the right to close.
A trustworthy Agent shouldn't just say "task completed." When evidence is insufficient, it should stop at unknown. Recoverable tasks continue; only when completion conditions are satisfied does it end. Define the endpoint clearly, keep the evidence.
References
OpenAI Agents SDK, Running agents (https://openai.github.io/openai-agents-python/running_agents/)
Malcolm Featonby, Making retries safe with idempotent APIs (https://aws.amazon.com/builders-library/making-retries-safe-with-idempotent-APIs/)
Stripe, Idempotent requests (https://docs.stripe.com/api/idempotent_requests)
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
Architect
Professional architect sharing high‑quality architecture insights. Topics include high‑availability, high‑performance, high‑stability architectures, big data, machine learning, Java, system and distributed architecture, AI, and practical large‑scale architecture case studies. Open to ideas‑driven architects who enjoy sharing and learning.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
