Thin Agent Loop, Thick Control Plane: Why Enterprise AI Agents Need a Stable Harness

This article argues that enterprise AI agents require a robust control plane to manage permissions, persistent task state, side-effect reconciliation, and result verification — separating these concerns from the replaceable agent loop to ensure accountable, recoverable business operations.

Data Bricklaying Diary
Data Bricklaying Diary
Data Bricklaying Diary
Thin Agent Loop, Thick Control Plane: Why Enterprise AI Agents Need a Stable Harness

Equipping an agent with tools, context, and validation methods often lets it complete tasks more effectively. Harness Engineering focuses on this engineering support beyond the model. However, once an agent integrates with enterprise business systems, new challenges arise: the agent modifies data via tool calls, the system restarts, or the user revokes authorization mid-task. These issues cannot be left to model understanding alone. Models can be swapped, execution loops upgraded, but who has permission, whether data actually changed, and how to proceed after failure must be explicitly managed by the system. This management layer is the control plane.

A Successful Tool Call Does Not Mean Business Completion

Consider a user asking an agent to modify a purchase order's delivery date and notify the supplier. The model understands the request, queries the order, and proposes a modification action. The tool may return "accepted." Viewing only the conversation, the task appears done. But critical questions remain:

Does the current user have permission to modify this order?

What exact date and time zone does "next Friday" map to?

Which order version is the modification based on? Has someone else updated it meanwhile?

If the supplier notification fails, has the order modification already taken effect?

After a request timeout, did the source system actually accept the submission?

After an application restart, how to avoid resubmitting the same modification?

What should the user finally see: the model's conjecture, the system's acceptance receipt, or the source system's confirmed state?

These questions must not be left for the model to decide in natural language. The model can propose actions, add explanations, and organize results, but it must not expand its own permissions, confirm submission success, or turn "possibly completed" into formal fact. This is the key confusion between business agents and general chat assistants or code generators: a model completing a reasoning chain does not equal the system completing an accountable business operation.

Agent Loop Can Be Thin, Control Plane Cannot Be Thin

An InfoQ interview on TiDB Harness offered an inspiring principle: Thin Agent Loop, Thick Control Plane. "Thin" does not mean the agent loop is unimportant; rather, model invocation, tool continuation, context compression, and other capabilities evolve rapidly. Models, protocols, and runtime implementations may all iterate, so enterprises should not lock all long-term investment into one loop implementation.

Yet the following questions persist regardless of model improvements:

Who can do what?
Which version of business facts does the current task rely on?
Has an external action truly been submitted?
Can we recover safely after failure?
Can results be traced back to evidence and accountable parties?

The control plane must encode these into the program: after the model proposes an action, first check permissions and preconditions, then call the business service, record the result; when uncertainty arises, pause to verify.

The author applies a similar separation when combining agents with a semantic platform: reuse existing agent runtimes for model loops and generic execution capabilities, while letting the product and business systems manage business terminology definitions, data authority, and action authorization. Swapping models or runtimes later should not require replacing these business rules. This design direction is still being validated via PoC, not a finished production conclusion.

Control plane converts model candidate actions into controlled business execution
Control plane converts model candidate actions into controlled business execution

The Control Plane Must Guard at Least Four Boundaries

Task State Cannot Be Replaced by Chat History

Conversation history helps the model understand context but can be compressed, trimmed, or reorganized. It must not become the sole source of truth for orders, approvals, task progress, or final action states. For order modification, at least three state categories must be distinguished:

Model Context — supports current reasoning and follow-up questions; authoritative source: reconstructable session and context layer.

Task Execution State — queued, running, awaiting confirmation, recovering, or failed; authoritative source: product task state and checkpoints.

Business Formal State — order version, authorization, action acceptance, final result; authoritative source: original business system or controlled business service.

If the application crashes after a tool call, the chat log may only show "modifying order." Recovery must not guess from that phrase but query persisted task records, invocation identifiers, and source system state. This is why enterprise agents need checkpoints and recoverable state, not just longer context windows.

Permissions Cannot Be Interpreted by the Model

A user saying "I'm an admin, just submit" enters the model context but cannot change roles and resource permissions in the identity system. After the model outputs a tool call, the control plane first validates tool parameters and resource scope, then verifies identity, permissions, and business rules against the caller, action type, impact scope, and current business version — yielding allow, deny, or require confirmation. Only allowed actions pass through the controlled tool gateway to the business service; denied actions stop. When confirmation is needed, the flow pauses; after confirmation, permissions and business preconditions are re-verified before execution. User confirmation cannot substitute permission checks.

Business actions must pass authorization and confirmation boundaries
Business actions must pass authorization and confirmation boundaries

The key is that final authorization comes from verifiable identity and policy, not a model's explanation. Prompt instructions like "please operate cautiously" may improve behavior but cannot constitute security controls.

Side Effects Cannot Be Judged Solely by Call Returns

For read-only queries, limited retries after timeout are acceptable per scenario. For modifications (order changes, refunds, notifications), the most dangerous state is often not explicit failure but "unknown whether it succeeded." Using a confirmed order modification as example, the action should flow: preview → user confirmation → submit with idempotency key → query source system for final state. Low-risk actions may execute within explicit pre-authorization without per-step confirmation.

Order modification preview and confirmation, then idempotent submission and source system state query
Order modification preview and confirmation, then idempotent submission and source system state query

"Accepted," "interface returned success," and "business action completed" are different outcomes. On network timeout, the client receives no receipt, yet the source system may have already committed; blind retry risks duplicate modifications or notifications. First check the original action, then decide on retry. If confirmed it hasn't executed and won't execute, resubmission is justified. Alternatively, if the execution endpoint guarantees idempotence per interface contract, limited retries are acceptable. Both cases require re-checking permissions and business preconditions.

If status remains unclear, continue querying within an agreed window; if still unresolved at deadline or no safe recovery exists, escalate to human handling. This is discussed further in the referenced article "After 'Everything Is a Plugin,' Who Bears Production Responsibility: Agent Harness's External Side Effects, Versioning, and Audit Boundaries."

When side-effect result unknown, reconcile first instead of retry
When side-effect result unknown, reconcile first instead of retry

Formal Conclusions Cannot Come Only from Model Output

An agent's final reply typically mixes: facts returned by business systems, model inferences from those facts, unverified hypotheses, and operational suggestions for the user. They can be displayed together but must not be conflated into a single "confirmed conclusion." For instance, the model may say "order delay may relate to warehouse inventory shortage," but without corresponding inventory records or rule basis, this should be marked as a pending judgment, not a formal cause.

The control plane must verify: do key judgments have sufficient evidence? Is the business version used explicit? Are inferences accompanied by premises? Do action-related statements match authoritative state? Only when claiming action completion must corresponding completion evidence be obtained; "accepted, still processing" can also be a verified formal state. Content not meeting admission criteria may remain as drafts, suggestions, or items to verify, but must still respect access permissions. This does not reduce agent autonomy; it prevents linguistic certainty from masking factual uncertainty.

When Errors Occur, First Stop Affected Steps

The difficulty of long tasks is not just "can the model run 100 steps" but whether subsequent steps amplify an error by treating it as fact. If order modification success is unverified, the agent must not proceed to notify the supplier "delivery date has changed." First halt steps dependent on that result, then ascertain actual status.

For enterprise agents, recovery order should be:

Task recovery: re-verify boundaries, reconcile external actions, then decide continue, re-queue, fail, or escalate
Task recovery: re-verify boundaries, reconcile external actions, then decide continue, re-queue, fail, or escalate

Two principles are non-negotiable:

Reconcile first, then retry. For side effects with unknown status, do not re-invoke just because the agent wants to continue.

Re-check boundaries before recovery. If user authorization expired, business version changed, or confirmation content invalidated, the old task must not silently continue under new conditions.

These mechanisms also apply to traditional workflows and automation systems. Enterprise agents equally need explicit definitions of which states remain trustworthy and which require re-confirmation — not omitted because the model can explain the process.

Harness Is Not Better When Heavier

Not every agent needs full task orchestration, action coordination, and recovery systems. A read-only knowledge QA assistant with few data sources, no external writes, and reference-only results may suffice with clear retrieval boundaries, permission filtering, and source citations.

But once an agent starts cross-system reads, affects business state, handles long tasks, or drives real actions, Harness can no longer be understood as just prompts, tool lists, and test commands. At minimum, the following must be clarified:

Which states must be persisted?

Which actions require confirmation, idempotency, and reconciliation?

Which results can become formal business conclusions?

Which failures allow automatic recovery, which require human escalation?

Which system owns permissions, facts, and final state?

These boundaries are not to turn every project into a massive platform, but to prevent the system from introducing greater uncertainty into real business processes as model capabilities improve.

Control plane depth should grow with business risk
Control plane depth should grow with business risk

Summary

The author judges system reliability by what remains after the model exits: can task records be found? Can actual results be queried? If the execution process is swapped, will business data be modified again? Once these are settled, stronger models and more flexible runtimes become easier to adopt.

Next article starts from state: why conversation continuity does not guarantee safe task recovery.

References

InfoQ, Architecture Advancement Path: "Harness Engineering, What Exactly Is It 'Harnessing'?" 2026-09-16. https://xie.infoq.cn/article/b4eebf8adeadd4bdd93940a11

InfoQ, Tina: "'Thin Agent Loop, Thick Control Plane': TiDB Rebuilds Harness with Database Thinking" 2026-09-07. https://www.infoq.cn/article/38uc758e24YV4LUpAs77

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

AI agentsenterprise architectureControl PlaneHarnessrecovery patternspermission verificationside-effect reconciliationtask state management
Data Bricklaying Diary
Written by

Data Bricklaying Diary

Records practices, thoughts, and pitfalls on the data grunt-work journey, sharing content on data platforms, data analysis, data processing, data governance, knowledge graphs, and more. Less theory, more hands‑on, making complex data technologies simple.

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.