DeepSeek Harness Architecture: Coordinating Multi-Turn LLM Calls, Tools & State

This article analyzes DeepSeek Harness's architecture, detailing its plugin assembly via Cordis, core agent loop with turn/step boundaries, append-only Session log with surface projection, tool pipeline with pre/post-execute hooks and concurrency control, sandbox-based security isolation, and exception handling via retry, compaction, and cancellation mechanisms.

Architecture and Beyond
Architecture and Beyond
Architecture and Beyond
DeepSeek Harness Architecture: Coordinating Multi-Turn LLM Calls, Tools & State

This analysis covers the DeepSeek Harness framework as of commit 477b4f4205 (2026-09-25). The framework addresses coordination between multi-turn model calls, tool execution, and session state: when input enters a request, how tool results are written back to history, where to retry after failures, what information remains after context compression, and what state can be recovered after interruption.

1. Plugin Assembly

DeepSeek Harness uses Cordis for plugin organization. Cordis provides Context, Service, Plugin, and Event abstractions; plugins collaborate via dependency injection and an event system, with lifecycles managed by the framework. Core components — LLM service, Session, Agent, AgentLoop, tool registry — are mounted via a bundle. At startup, the dsh CLI loads configuration by profile, locates config lines by ID, and replaces the entire config segment ( whole-segment replacement, not field-level merge ). Overwriting a component's config may drop fields not explicitly retained.

This design enables component swapping and capability composition, but spreads execution logic across multiple plugins. Analyzing a single request requires examining the core loop, plugin registration, and the final effective configuration simultaneously.

2. Core Loop

The agent-loop drives the model-tool cycle. Its main execution path:

Prepare model request, parse config, select adapter.

Project system prompt, submit user input that must enter history.

Build request from session, receive and accumulate streaming response.

Write assistant message to Session.

Execute model-proposed tool calls, feed results to next model request.

End turn when no tool calls or termination condition met.

Three boundary concepts clarify this flow:

turn : user-input-driven round.

step : execution boundary within a turn.

request attempt : a single actual model call; retries stay within the same step.

A step ≠ one network request. Retries may produce multiple attempts but must not re-consume input or re-submit user messages. The agent maintains next-turn and next-step queues for pending inputs. Queue modifications are recorded as agent/inbox/spliced events in the Session, allowing recovery from the log. This separates "system received message" from "model consumed message". Incoming supplemental requirements enter at the appropriate boundary; they cannot alter an already-sent model request or undo executed tool side effects. Logging queue changes prevents losing unprocessed inputs during recovery.

3. Session

Session implements an append-only event log with incrementing sequence numbers, type for event classification, and body for payload. It records not only user/assistant messages but also request failures, queue changes, tool executions, and compression processes.

Crucially, events in the log do not automatically enter the model context. The actual messages sent to the model are rebuilt by the Session's surface via deriveMessages() and frozen into an immutable request. Only events that enter the surface participate in history projection; runtime records like assistant/attempt do not become model messages.

This separation serves two purposes:

Preserve debugging info : failures and recovery are fully logged without consuming context window.

Allow adjustable visible history : compression or replacement of model context retains original events.

Appending to Session validates JSON serializability, event rules, and surface replacement legality to reduce unpersistable or incorrectly projected data.

Log recovery ≠ external action rollback. Append-only logs provide audit and state reconstruction but cannot guarantee lossless recovery at arbitrary interruption points. Example: a tool modifies a file but the process exits before writing the result to Session — external state and log diverge. Recovery can restore committed records but cannot infer from "missing result" that the tool didn't run. Thus recovery splits into:

Restoration of committed events and their projections.

Confirmation and handling of external side effects (requiring tool idempotency, state checks, or extra transaction mechanisms).

4. Tool Pipeline

The tool pipeline moves from model intent to controlled execution. The registry exposes only name, description, parameters to the model; execution function, timeout, and concurrency attributes stay at runtime. This separates model-side description from runtime execution control: the model proposes calls, the framework decides executability and execution method.

A tool call passes through:

Parse model-generated parameters.

Find tool visible to current agent.

Run tools/pre-execute checks (allow/deny).

Enter tools/execute wrapper chain.

Process result via tools/post-execute.

Write result back to Session for subsequent model requests.

Permission checks, execution wrapping, and result handling thus have independent intervention points.

Concurrent execution, sequential commit. Multiple tool calls from one assistant message are scheduled by the tool's isConcurrencySafe flag:

Explicit true → parallel pool.

Others → exclusive barrier.

Concurrency is opt-in; undeclared tools never run in parallel by default. Tool bodies may run concurrently, but preparation, post-processing, tool/result writes, and additional context enqueueing are committed in model-call order. This prevents random completion times from changing result order in the session. Trade-off: possible waiting — a later tool that finishes early must wait for earlier calls to commit. Concurrency shortens execution time but does not guarantee immediate history entry; early results may be buffered. Moreover, stable commit order ≠ deterministic external side effects ; actual safety depends on shared-state access, execution dependencies, and accuracy of the concurrency declaration.

5. Security Isolation

DSH enforces approval and execution restrictions at separate layers for security isolation. Sandbox modes:

type SandboxMode =
  | 'read-only'
  | 'workspace-write'
  | 'danger-full-access'

Policies cover file-path permissions, command execution limits, and privilege escalation. Tools can request elevated rights via justification and sandbox_permissions, processed by an approval mechanism. The base bundle mounts a workspace-write file sandbox and an ask approval policy; sandbox-policy parses the current file-operation mode for the actual backend.

Three distinct layers:

Tool visibility : which calls the model can propose.

Execution approval : whether a specific call gets permission.

Backend isolation : what resources the approved execution can actually access.

These layers are not interchangeable. Approval only means the action is permitted, not that the environment has unlimited rights; a tool hidden from the model doesn't prove other tools cannot reach the same resource. The effective security boundary depends on both policy configuration and the concrete backend implementation, not merely the sandbox mode name.

6. Exception Convergence

DSH handles exceptions via retry, compaction, and cancellation.

6.1 Request Retry

On model request failure, the system first writes assistant/attempt, then passes to the agent/request-error waterfall chain. llm-retry uses the adapter's retry policy to classify error, count retries, and determine backoff. Planned retries and actual starts are logged as llm/retry and llm/retry-started. Retries re-prepare the request within the same step, avoiding re-submission of the initial user message — this boundary prevents network-layer retries from becoming session-layer duplicate inputs. Separate logging of plan vs. start lets the log distinguish "decided to retry" from "started next attempt".

6.2 Context Compaction

Auto-compaction triggers at agent/pre-step when context pressure is detected, logging compaction/start, compaction/summary, compaction/end. The system then replaces a selected contiguous history interval with a single checkpoint user message. This is a surface replace : model-visible history shrinks, but the original log remains. Compaction reduces request context, not persistent storage. It introduces potential information loss: original constraints still in the log but absent from the summary or remaining surface become unavailable to subsequent model turns. Effectiveness depends on whether task goals, constraints, and unfinished items are preserved, not just context length.

6.3 Cancellation

AbortSignal

propagates through turns, requests, and tools, providing a unified stop-signal path. However, signal propagation ≠ immediate halt of all actions. Tool exit timeliness depends on whether it checks and responds to cancellation; completed file writes or remote operations are not auto-rolled-back by AbortSignal. Cancellation supplies a control path to stop further execution; external side-effect reversal requires independent implementation.

7. Summary of Trade-offs

DeepSeek Harness makes explicit architectural trade-offs between traceability and runtime cost. Key mechanisms and their boundaries:

Cordis plugin assembly — enables component replacement and capability composition; cost: behavior distributed across config and multiple handlers.

turn/step with input queues — clarifies input ownership and execution boundaries; cost: added queue state and event maintenance.

Append-only Session — provides audit and committed-state reconstruction; cost: log grows unbounded, does not guarantee exactly-once external actions.

Surface projection — separates runtime log from model context; cost: requires maintaining projection consistency and replacement rules.

Concurrent execution, sequential commit — leverages parallelism while stabilizing history order; cost: possible commit waiting and result buffering.

Sandbox and approval — layered action-permission control; cost: actual guarantees depend on config and execution backend.

Auto compaction — reduces context-window pressure; cost: adds summarization overhead and may lose task information.

AbortSignal — unified cancellation intent propagation; cost: relies on tool cooperation, cannot auto-rollback side effects.

The core value of DeepSeek Harness is making critical Agent execution states explicit: inputs have queues, execution has boundaries, requests have visible history, tools have scheduling and approval, failures have dedicated handling paths. These mechanisms enable multi-turn tasks to be traced, inspected, and recovered, while clearly delineating framework limits: log recovery cannot substitute side-effect confirmation, sequential commit cannot replace concurrency safety, approval cannot replace environment isolation, and context compaction cannot guarantee lossless information retention.

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

Agent FrameworkSession ManagementLLM OrchestrationSandbox SecurityContext CompactionCordisDeepSeek HarnessTool Pipeline
Architecture and Beyond
Written by

Architecture and Beyond

Focused on AIGC SaaS technical architecture and tech team management, sharing insights on architecture, development efficiency, team leadership, startup technology choices, large‑scale website design, and high‑performance, highly‑available, scalable solutions.

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.