AI Agent Architecture: The Runtime Chain Connecting Planner, Tool, MCP, Memory & Reflection
This article deconstructs AI agent architecture as a stateful runtime chain—Goal→Plan→Context→Decide→Act→Observe→Verify→State→Next—distinguishing Planner, Tool, MCP, Memory, and Reflection as interconnected layers rather than parallel modules, emphasizing that reliability comes from the harness runtime handling verification, recovery, and context engineering outside the model.
TL;DR
Agent is not a single LLM call. A call only handles input and output; an agent must manage goals, tools, state, verification, and recovery.
Planner, Tool, MCP, Memory, Reflection are not on the same layer. Some govern decisions, some connections, some state, some feedback.
Function Calling is the invocation expression; MCP is the protocol layer. Tool is the capability itself; Harness wires permissions, logs, retries, and verification.
Memory ≠ chat history. Context window, event log, and reusable experience must be separated.
Reflection is not model self-praise or self-criticism. It must connect to tests, metrics, human confirmation, rules, and next-round context.
Architectural reading returns to the execution site. First see how the goal enters, then how actions happen, finally how evidence is retained.
Figure 1: An agent is not a single call but a runtime chain with state, verification, and recovery.
Agent Is Not a Single LLM Call
A single LLM call has the structure "input context, output text." The model can explain problems, generate plans, produce code, or emit structured JSON. After the call ends, the system does not automatically know which file to read next, which test passed, or which API call already produced side effects.
An agent adds a software structure that lets the model continuously advance within a task instead of answering once. For example, a user says: "Help me troubleshoot this API timeout, fix the issue, add tests, and summarize changes." A Q&A model only gives troubleshooting suggestions. An agent treats the sentence as a task: understand the goal and acceptance criteria, then decide to read logs, inspect traces, examine code, run tests, modify files, review diffs, and finally deliver the summary.
Anthropic's Building effective agents distinguishes workflows from agents. A workflow runs LLMs and tools along a pre-written code path; an agent lets the LLM dynamically decide the flow and tool usage based on current state. This affects architectural trade-offs: stable paths with clear branches and acceptance rules suit workflows (low cost, clear observability). Tasks with unknown branches that require the system to judge the next step while executing call for an agent.
Discussing agent architecture is not just about "whether there are seven modules," but about how much dynamism the task needs, which parts are coded into fixed paths, and which judgments are left to the model.
Workflow vs Agent Comparison
Flow path: Workflow — pre-written by developers; Agent — model chooses next step based on current state
Input variation: Workflow — fields, formats, rules relatively stable; Agent — unknown branches appear mid-task
Tool usage: Workflow — fixed sequence or few branches; Agent — dynamically chosen based on observations
Acceptance: Workflow — explicit rules, fixed checks; Agent — requires phased evidence and human confirmation
Suitable tasks: Workflow — review, extraction, reporting, fixed tickets; Agent — troubleshooting, research, code modification, long-running execution
A practical heuristic: If you can write a clear state machine, do it as a workflow first; only hand the uncertain parts to an agent when the state machine is incomplete.
Goal: The Goal Must Become Acceptance Criteria
The entry point of an agent is not a pretty prompt but a goal to be completed. This step looks simple but influences downstream boundaries. If a user says "do an industry research," the system must clarify the deliverable, scope, sources, whether to download originals, whether to provide judgments, and what standard marks completion. Ambiguity here skews every subsequent module: the Planner produces wrong steps, Tools query wrong places, Memory retains irrelevant information, and Reflection only self-checks against the wrong goal.
Architecturally, Goal is not the raw user input; it is closer to a task contract:
Deliverable: what form the final output takes — answer, document, code, ticket, or executable change.
Constraints: usable sources, forbidden operations, time, cost, permission limits.
Acceptance: what evidence proves completion, what situations must return to human.
Risk: which actions have side effects, which inputs are sensitive.
The contract need not be heavy; simple tasks need a few sentences, complex tasks need explicit boundaries. Downstream modules must know which result they are working toward.
Planner: Not Writing All Steps at Once
Planner is often described as "breaking a large task into small steps." That is correct but implies planning is a one-time roadmap generated at the start and then followed. Engineering tasks rarely follow the initial plan to the end. An API timeout may stem from slow SQL or an external service retry; document writing may discover halfway that the source version changed; code repair may expose another edge case during testing. A plan that cannot adjust with observations fails early.
Here, Planner manages phases, not a one-shot script. It decides the current phase's goal, allowed tools, required evidence, and conditions to enter the next phase. More complex systems make the plan replayable steps, each with input, action, output, and state.
Planner must be separated from Workflow. Workflow is a developer-defined process, suitable for stable tasks like fixed-format review, field extraction, standardized reports. Planner handles dynamic tasks: choosing the next step given current state, reordering phases when needed, even narrowing the goal or asking humans for more information. The two are not mutually exclusive; mature systems mix them: stable parts solidified as workflows, unknown parts left to agent planning. This approach is not new but keeps control costs low.
Context: The Agent's Next Step Is Decided by Context First
Some agent architecture diagrams place Reasoning in the middle, as if the model's thinking ability determines subsequent results. The model matters, but what judgment it can make is first limited by what it sees in this round's context.
Karpathy's "context engineering" replaces "prompt engineering" — not writing prettier prompts, but placing task instructions, examples, RAG results, tools, state, history, compressed summaries, etc., into the next-step context in appropriate forms. Insufficient information gives the model no basis; overload raises cost and noise. Another trouble: some information must not enter the model window directly. Tool call logs, permission records, full file diffs, external API responses must enter system records, but the next round only needs a small subset.
The Context layer must first separate three information types:
Information Types in Context Layer
Current Context — Primary use: support next-step judgment; Where stored: model window; Sent to model every round: yes, but filtered
Event Log — Primary use: audit, recovery, replay; Where stored: system storage; Sent to model every round: not necessarily
Reusable Memory — Primary use: avoid detours in future tasks; Where stored: memory library or rule library; Sent to model every round: recalled on demand
Current goal, phase plan, available tools, key observations, and recent failure reasons can enter the model context. Full tool parameters, raw returns, approval records, and execution latency belong in the event log. User preferences, project rules, and stable failure patterns must be validated before entering long-term memory.
Equating Memory directly with chat history narrows system boundaries. Chat history is only historical material. An agent needs retrievable, compressible, auditable, migratable state.
Reasoning: Responsible for Local Decisions, Not Whole-System Reliability
The ReAct paper places reasoning and acting in the same loop. Its key point is not making the model "think longer," but letting the model produce a reasoning trace while executing task-relevant actions; actions connect the model to external knowledge or environment, and observations feed back into the next step.
In engineering, Reasoning corresponds to the current round's local decision. The model sees the goal, context, plan, tool descriptions, and previous result, then judges what to do next — query a log, read a config, or stop and deliver.
But Reasoning does not cover three responsibilities:
It cannot enforce permissions. The model saying "delete a file" does not mean the system can delete it.
It cannot persist state. The model can recount history, but that does not give the system a recoverable event record.
It cannot perform acceptance. The model saying "fixed" cannot replace tests, metrics, code diffs, or human confirmation.
Some implementations stuff these responsibilities into the prompt: remind the model to be careful with tools, to self-check, to remember rules. Prompts help, but they are not architectural boundaries. Permissions, state, verification, and recovery must live in the runtime outside the model.
Tool: The Model's Action Space
Tool is how an agent affects the external world. Search, read/write files, query databases, send HTTP requests, run tests, open browsers, create tickets, send messages — all are tools. Without tools, the model only outputs text; with tools, the model gains an action space.
But more tools are not always better. More tools enlarge the model's choice space, increasing invocation cost, permission risk, and debugging difficulty. A search tool and a production database write tool carry vastly different risks. An idempotent query and a refund action do not belong in the same confirmation strategy.
The tool layer must answer:
What is the input schema; can parameters be strictly validated?
Does the tool have side effects; can they be undone or compensated?
Is a permission check or human confirmation required before invocation?
How are errors returned to the model on failure?
How large is the result; what enters context, what stays only in logs?
Does the tool have timeout, cancellation, retry, and idempotency design?
When integrating tools, I first layer by side effect:
Tool Classification by Side Effect
Read-only query — Examples: search, read file, query log; Primary risk: result pollutes context; Architectural handling: summarization, citation, source tagging
Mutable local state — Examples: write file, change config, run script; Primary risk: overwrite, accidental delete, irreproducible; Architectural handling: diff, backup, approval, rollback
External system impact — Examples: send email, refund, modify order; Primary risk: real side effects; Architectural handling: permission, confirmation, idempotency key, audit
High-cost invocation — Examples: large-scale scraping, long-running tasks; Primary risk: cost and resource runaway; Architectural handling: quota, timeout, cancellation, budget
OpenAI Agents SDK wraps function tools with schema and Pydantic validation, and provides guardrails, sessions, tracing primitives. This combination shows: a tool is not an isolated module. Once a tool enters the Agent Loop, it pulls in input validation, state persistence, flow tracing, and safety checks.
Function Calling: Turning Intent into Structured Invocation
Function Calling is often conflated with Tool, but Function Calling is the "invocation expression method." The model does not execute functions directly; it outputs "which function I want to call and what the parameters are." The host application parses this structure, validates parameters, then decides whether to execute the corresponding tool.
A boundary must be clear here. If the model outputs natural language like "help me query the database," the host must guess. Function Calling turns that into a structured object, e.g.:
{
"name": "query_orders",
"arguments": {
"user_id": "u_123",
"limit": 5
}
}After structuring, the system can perform schema validation, permission judgment, audit logging, and error handling. Invalid parameters are rejected; insufficient permissions trigger confirmation; tool timeouts return explicit errors.
Function Calling solves "how the model's intent becomes machine-readable." It does not automatically solve whether the tool is safe, whether the result is correct, or whether the action should happen.
MCP: Protocol Layer, Not the Tool Itself
When tool count grows, another problem appears: each system integrates a separate API, turning the agent host into an integration quagmire. MCP (Model Context Protocol) solves the connection method. The 2025-06-18 spec defines Host, Client, and Server roles. Host is the LLM application initiating connections; Client is the connector inside the host; Server provides context and capabilities.
MCP Server offers three capabilities: Resources, Prompts, Tools. Client side also has Sampling, Roots, Elicitation. From this design, MCP is not just "standardized tool list." It also brings data, prompt templates, user-supplemented information, and file boundaries into the protocol discussion.
However, MCP is not the complete security boundary. The spec explicitly mentions user consent, data privacy, tool safety, and sampling control, but also states the protocol itself cannot enforce these principles at the protocol layer. MCP standardizes message and capability exposure, but authorization, audit, sandboxing, data desensitization, credential isolation must still be implemented by the host application and organizational policies.
This clarifies the relationship among Tool, Function Calling, and MCP:
Tool is the capability itself (search, query DB, write file).
Function Calling is the model's structured way to express invocation intent.
MCP is the protocol layer that connects external context and capabilities into the agent host.
Harness is responsible for placing the invocation into permission, state, verification, and recovery mechanisms.
Figure 2: Model expresses invocation intent; host handles validation and governance; MCP connects external capabilities.
Separating these layers reduces confusion when discussing tool integration.
Observe: Tool Results Are Not Directly Stuffed Back into the Model
After a tool executes, it returns an observation result. This looks like an intermediate product in the Agent Loop, but problems often arise here.
Result volume may exceed what the model needs next round. Search results, logs, test outputs, web pages, database queries — stuffing all back into context wastes tokens and interferes with judgment. Result sources may be untrustworthy: web pages, user-uploaded files, third-party system responses may contain content that induces the model to change rules. The system must distinguish "data returned by the tool" from "content that can be treated as instruction." Some results carry side-effect evidence: an order already refunded, a file already modified, an email already sent. These cannot rely on the model to remember; they must enter the event log.
The Observe layer does two things: save the raw result, then put a model-readable summary, evidence, or error information back into context. Part of system stability difference lies here — not a stronger model, but tool results that have been filtered, summarized, cited, and error-categorized.
Memory: Separate Context, Events, and Reusable Experience
Memory is a concept with wide boundaries in agents. A common taxonomy is short-term, long-term, working memory. This helps, but engineering needs one more layer: where exactly they live, when they are read, and in what form they enter the next round.
In implementation, Memory can be split into three objects:
Memory Objects
Context Window — Facing: model's next decision; Stores: current goal, plan, key observations, failure reasons; Common mistake: stuffing full logs into it
Event Log — Facing: system recovery and audit; Stores: tool parameters, returns, latency, errors, approvals, side effects; Common mistake: letting the model "remember" only
Reusable Memory — Facing: future tasks; Stores: user preferences, project rules, stable failure patterns, validated experience; Common mistake: turning a one-off result into a permanent rule
Figure 3: Context window, event log, and reusable memory solve three different problem classes.
OpenAI Agents SDK calls Sessions the persistent memory layer that maintains the Agent Loop's working context, and provides Tracing to visualize, debug, and monitor workflows. This gives a reference: Session keeps tasks continuous, Tracing makes the process visible, Memory makes experience reusable.
These three layers cannot substitute each other. Chat history helps the model review conversation, but task failure recovery relies on event logs. If event logs don't enter context assembly, the model still misses the focus next step. Long-term memory also needs source validation; otherwise the system may turn a one-off failure into a permanent rule.
The key to Memory is not "store more," but "when to take what out."
Reflection: Reflection Must Connect to Verification and Improvement
Reflection is often translated as "反思" (self-reflection). This translation shifts focus to model self-review, as if next performance would naturally improve. The Reflexion paper provides an important direction: it does not update model weights, but lets the agent generate verbal reflections based on task feedback and places them in an episodic memory buffer for subsequent trials. The value: agent improvement happens not only in model weight training, but also in language feedback, memory, and next-round decisions.
In production systems, Reflection cannot rely solely on model self-feeling. Feedback sources differ by task: code tasks look at unit tests, type checks, browser screenshots, logs, code reviews; support tasks look at order status, refund results, user confirmation; writing tasks check original text verification, citation checks, structure review. The model interprets failures, summarizes causes, proposes next steps. But "whether it passes" should be judged by verifiable validators.
In the system, Reflection can be split into two layers:
Feedback generation: extract problems from test failures, tool errors, user supplements, review comments.
Feedback effectiveness: write stable problems into next-round context, rules, eval samples, skill documents, or versioned configs.
Figure 4: Reflection's value is not in summarization, but in bringing failure causes into rules, samples, and next-round context.
JGX's recent articles point to this direction. The Sep 25 piece on Harness discussed the model proposing the next step while Harness puts it into a controllable runtime. The Oct 3 piece on Claude.dev emphasized that errors must leave evidence, entering eval, rules, and versioned improvements. The Oct 4 piece on Claude Code Mods showed runtime control points being productized. These three are not scattered topics; placed in agent architecture, they all point to the same question: after an agent errs, how does the system make that error useful for the next time?
Output and Interaction: Delivery Is Not Just the Last Sentence
Agent output also deserves separate examination. In demos, the agent ends with "task completed." But real systems deliver PRs, reports, spreadsheets, email drafts, support tickets, database changes, deployment records, or evidence-backed conclusions. Here, Output is not just a natural language answer, but a combination of task status, artifacts, and evidence.
Interaction changes accordingly. Ordinary Chat UI suits Q&A, but long-running agent tasks need users to see the plan, current phase, active tool calls, produced changes, risk confirmations, and rollback points. Diffs, test results, run logs, and approval buttons in developer tools are all part of agent interaction.
Andrew Ng mentions three loops: agentic coding loop, developer feedback loop, external feedback loop. An agent can write code, test, and iterate in minutes; developer feedback may appear after tens of minutes or hours; external user feedback cycles are longer, waiting for canary, A/B tests, or real usage. These three loops operate on different time scales and cannot all be squeezed into the model's single thinking process. The Agent Loop solves the immediate next step; developer feedback decides goals and trade-offs; external feedback decides whether the system truly improves long-term. Architecture that doesn't separate these loops mistakes "model keeps going" for "system is improving."
Harness: Wiring Modules into a Controllable Runtime
The previous layers must work together, requiring a Harness outside the model to wire them. Harness is not necessarily a standalone product; it is a set of runtime responsibilities:
Context assembly: decide what the model sees next round.
Tool routing: turn model intent into executable actions.
Permission control: judge whether an action is allowed, needs confirmation.
State persistence: record plans, events, checkpoints, side effects.
Verification mechanism: use tests, rules, metrics, or human confirmation to judge results.
Tracing and observability: make the process debuggable, reviewable, replayable.
Recovery and stop: handle timeout, cancellation, retry, rollback, human takeover.
Harrison Chase's retelling of the Meta-Harness discussion notes that an agent's continuous learning happens not only at the model layer, but also at the harness layer and context layer. This matches engineering practice. System improvements sometimes come not from model weights, but from context assembly, tool definitions, validation rules, eval samples, and recovery strategies. That is why Harness is pulled out separately here.
If you only watch model output, you treat the agent as a "thinking interface." If you only watch the tool list, you treat it as an "API-calling robot." Following the Harness view, an agent is a software system that puts uncertain decisions into a controlled process.
Five Concept Pairs Separated
Finally, collect a few concept pairs:
Concept Pair Clarifications
Planner = Workflow → More accurate split: Planner handles dynamic phases; Workflow solidifies stable paths. Architectural impact: stable parts in code; unknown branches to model judgment.
Reasoning = Reflection → More accurate split: Reasoning faces current round; Reflection faces feedback improvement. Architectural impact: current decision and long-term improvement designed separately.
Tool = MCP → More accurate split: Tool is capability; MCP is access protocol. Architectural impact: protocol cannot replace permissions, audit, sandbox.
Memory = Chat History → More accurate split: Memory includes context, event log, reusable experience. Architectural impact: recovery, audit, reuse need separate storage.
Agent = LLM + Prompt → More accurate split: Agent also needs loop, tools, state, verification, interaction. Architectural impact: difficulty often lies in the runtime outside the model.
In one sentence: Module names are just entry points; what truly affects system quality is how modules pass information, execute actions, and leave evidence.
From Architecture Diagram Back to Concrete Questions
When evaluating an agent architecture, I first look at what the model actually reads next round. Task instructions, history, tools, state, retrieval results, compressed summaries — which enter the window, which stay in the event log — the system must explain clearly.
Then examine tool invocation impact scope. Tool, Function Calling, MCP only describe how capabilities are expressed and connected; they cannot replace access control. Search, query DB, write file, refund, send message should have different authorization and confirmation strategies.
Failure recovery must be examined separately. Whether the system can find the last confirmed-complete action depends on checkpoints, idempotency design, and replayable logs. Without these records, retry becomes asking the model to guess again.
Finally, look at completion evidence. Test passes, metric recovery, file diffs, external interface states, human confirmations must bind to the task goal, not just the model's final "done."
These questions bring discussion back to the engineering site, not stopping at "has Planner, has Memory."
Evaluation Points and Questions
Context — Questions: what exactly does the model read next round? Evidence form: prompt fragments, retrieval results, summarization strategy
Tool Invocation — Questions: where are action boundaries and side effects? Evidence form: schema, permissions, approvals, idempotency keys
Failure Recovery — Questions: can the system return to a known-good state? Evidence form: checkpoints, event logs, replayable records
Completion Judgment — Questions: how does the task count as finished? Evidence form: tests, metrics, diffs, human confirmation
Long-term Improvement — Questions: do errors feed into the next run? Evidence form: rules, evals, Skills, long-term memory
Modules matter. Without Planner, complex tasks lack phase constraints; without Tools, the model can only talk; without Memory, long tasks lack continuity; without Reflection, errors don't enter the next round. But module names alone don't guarantee reliability. Reliability comes from the runtime relationships between modules: how information flows, how actions execute, how state persists, how results are verified.
Conclusion
Agent architecture can start with seven or eight modules, but should not stop at a module checklist. A practical view treats the agent as an end-to-end chain: goal enters the system, plan constrains phases, context feeds the model, model makes local decisions, tools and protocols connect to the external world, observations return to the system, validators confirm progress, state and memory support recovery and improvement.
When this chain runs through, Planner, Tool, MCP, Memory, Reflection become not just concepts but clear responsibilities within a single runtime. That is where the difficulty of agents lies. The key is not wrapping the model into a "thinking person." The more worthy investment is putting every uncertain next step of the model into a controllable, verifiable, recoverable system.
References
Anthropic, Building effective agents (https://www.anthropic.com/engineering/building-effective-agents)
Model Context Protocol, Specification 2025-06-18 (https://modelcontextprotocol.io/specification/2025-06-18)
OpenAI, OpenAI Agents SDK (https://openai.github.io/openai-agents-python/)
Shunyu Yao et al., ReAct: Synergizing Reasoning and Acting in Language Models (https://arxiv.org/abs/2210.03629)
Noah Shinn et al., Reflexion: Language Agents with Verbal Reinforcement Learning (https://arxiv.org/abs/2303.11366)
Andrej Karpathy, Context engineering (https://x.com/karpathy/status/1937902205765607626)
Andrew Ng, Loop engineering & three feedback loops (https://x.com/AndrewYNg/status/2071988145667928442)
Harrison Chase, Meta-Harness & harness layer learning (https://x.com/hwchase17/status/2040471961206214864)
Simon Willison, Context engineering (https://simonwillison.net/2025/Jun/27/)
From One LLM Call to Complete Harness: System Understanding of How Agents Work (2026-09-25)
Claude.dev's Four Engineering Articles: When Does an Agent Count as Truly Improved (2026-10-03)
Claude Code Mods Getting Started: Building a Mod from Zero (2026-10-04)
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
Architect
Professional architect sharing high‑quality architecture insights. Topics include high‑availability, high‑performance, high‑stability architectures, big data, machine learning, Java, system and distributed architecture, AI, and practical large‑scale architecture case studies. Open to ideas‑driven architects who enjoy sharing and learning.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
