DeepSeek Harness: Plugin-First Agent Architecture vs Claude Code

This article dissects DeepSeek Harness, an MIT-licensed agent framework where everything is a plugin, explaining its Cordis-based spatiotemporal composability, internal execution loop, and trade-offs against Claude Code's integrated approach.

Tencent Architect
Tencent Architect
Tencent Architect
DeepSeek Harness: Plugin-First Agent Architecture vs Claude Code

Introduction

On August 13, 2026, DeepSeek released Harness v0.1 developer preview under MIT license. The GitHub repository gained 50,000 stars in 12 hours, and 288 third-party plugins tagged dsh-plugin appeared within 24 hours. Harness is not a new model but an engineering framework that lets models "do work" — handling tool calling, task planning, execution scheduling, sandboxing, storage, and loops. It positions directly against Claude Code but with a fundamental difference: Claude Code is a finished product, while Harness is a developer substrate for building custom agent stacks.

What Is Harness

DeepSeek defines the equation: Model + Harness = Agent . The model handles reasoning; Harness manages everything else. Senior researcher Chen Delin stated: "Simply put, it benchmarks against Claude Code, building DeepSeek Code Harness." The core design principle is "Everything is a Plugin" — model, tools, skills, sessions, sandbox, storage, scheduling, and even the Web UI are replaceable, recomposable plugins.

Why "Everything Is a Plugin": Four Layers

Layer 1: Eliminate Lock-in

Components should not hold the system hostage. Claude Code is deeply tuned for the Claude model; if the framework is welded to a model generation, it ages with the model. Harness treats the LLM as a swappable part, supporting nearly 40 providers including OpenAI/Anthropic-compatible endpoints and Ollama local models. Similarly, swapping sandbox, storage, or tools in Claude Code requires peripheral extensions like MCP Hooks; Harness allows replacement at the configuration layer without touching source code.

Layer 2: Cordis Paper's Spatiotemporal Composability

The theoretical foundation is the Cordis framework and the paper A Programming Paradigm for Spatiotemporal Composability (Peking University + DeepSeek-AI, 2026-08-13). It formalizes dynamic runtime composition with two orthogonal guarantees:

Temporal Composability — Removable : When a component is removed, all its modifications to the shared environment must be fully and safely reverted, returning the environment to its pre-installation state. The paper cites a counterexample: 87 of the top 100 VS Code Marketplace extensions contain executable code and require a full host restart on uninstall/disable because their activation mutates shared state without symmetric cleanup. Cordis solves this with revertible effects : every state change carries an explicit inverse operation, tracked and composed at runtime; uninstall rolls back in reverse order. The pattern is ctx.effect(() => { ...; return dispose }) where dispose is the inverse.

Spatial Composability — Connectable : Components can structurally declare, discover, and resolve dependencies; when dependencies change, lifecycles coordinate automatically. This maps to reactive coeffects : a component declares ctx.llm, ctx.tools, and when the depended service is replaced or revoked, the dependent rebuilds automatically instead of holding stale references.

Both dimensions are necessary: temporal alone leaves undeclared dependencies and ordering errors; spatial alone leaves untracked global state on uninstall.

Layer 3: Target Is Self-Evolving Agents

The paper's true target is a self-evolving agent harness . Imagine an agent continuously processing requests while AI generates and deploys modifications to its own components — swapping tool implementations, adding/removing capabilities, replacing orchestration logic. Each mount/unmount is a dynamic composition. Without temporal composability, every self-modification would require a full process restart, discarding connections, caches, and in-flight turns; worse, a defective self-modification could crash the recovery process itself. Harness implements this via cordis_define / cordis_run tool families, letting an agent write, activate, and tear down a plugin mid-session.

This design is battle-tested: Cordis has run in the Koishi chatbot framework for four years with over 4,000 community plugins in production. Harness is a new application on top of Cordis.

Layer 4: Competing for "Evolution Rights" in the Agent Era

Strategically, DeepSeek uses MIT licensing and radical architecture to mobilize global developers while model capabilities and the Harness paradigm are still unsettled. Open weights solve "who can run the model"; open Harness tackles "who decides how the model works". DeepSeek aims to own the coordinate system in which future agent answers emerge.

Inside a Task Execution: Four Acts

To avoid jargon (profile, bundle, inbox, turn), the article walks through a concrete task: dsh --profile web then "Read README, understand project vision, list modules to design."

Act 1: Startup — Assembling the "Body"

Launch decides which plugins to load and in what order. Two concepts organize the assembly:

Profile : A named configuration list (e.g., web for web UI, headless for no UI).

Bundle : A pre-packaged plugin group (e.g., dsh-base for model adapter + tools + storage + sandbox; dsh-web-app for web UI).

They layer in fixed order: base bundle → profile bundles → user patches (one-line overrides). The result is a plugin tree rooted in the kernel. dsh --profile web --dump-config visualizes this tree.

Act 2: Receiving — How Your Instruction Is Caught

The instruction enters an inbox (queue). The agent-loop (scheduler) claims it, starting a turn — a complete segment from user input to final answer. A turn contains multiple steps (each: model thinks once + calls tool once). The scheduler prepares two packages for the model: (1) system prompt (identity, rules, environment) and (2) tool manifest (available tools with parameter schemas: read_file, edit_file, shell, web_search, …).

Act 3: Working — Think, Act, Observe, Repeat

Step-by-step walkthrough:

Model thinks, decides "read README first" : Scheduler sends system prompt + tool manifest + user instruction. Model outputs a tool call: read_file with parameter README.md.

Runtime actually reads the file via a three-stage tool execution pipeline:

Pre-execute approval check (sensitive ops like delete or shell prompt confirmation — Harness's approval policy ).

Execute: filesystem plugin reads README.md.

Post-execute: wraps result as a tool-result message.

Result returns to model, continues thinking : README content fed back. Model now wants to inspect src/ structure, emits shell (e.g., ls).

Loop repeats until model judges information sufficient.

Model submits final answer : Outputs complete module list. Scheduler sees no pending tool calls, ends turn.

Key insight: Agent work is not "one-shot answer" but incremental query-look-think cycles.

Act 4: After Submission — Why Every Step Is Replayable

Every event — system prompt, model inputs, model outputs, tool calls, results — is appended to an append-only session log with two invariants: (1) immutable history; (2) "model-visible implies recorded" enforced by runtime. This enables a Trajectory view to rewind, fork, and replay like git history. Example: if a decision before README reading was flawed, fork from that point, change prompt, re-run without restarting.

Three "Easter Eggs" Demonstrating Plugin Philosophy

Capability Seams : The file-system provider plugs into a seam ; swap ctx.fs and ctx.subprocess to remote sandbox implementations — file reads, Bash, LSP all migrate without forking code.

Runtime Self-Extension : Agent holds cordis_* tools ( cordis_define, cordis_run, cordis_stop …) to write, install, and remove plugins mid-run — the engineering landing of self-evolution. Dynamic code installation triggers approval.

Four Modes : Same instruction runs under four plugin combinations:

Standard: full toolset for daily development.

PTC (Programmatic Tool Calling): model generates code that orchestrates multi-step tool calls; suited for sequential batch ops.

Minimalist: only shell + file edit; for minimal-environment model benchmarking.

Creative: inspect runtime, experiment with Cordis plugins, compose new modes; for R&D.

Harness vs Claude Code: Wins and Losses

Premise: not "good vs bad architecture" but trade-off between freedom and out-of-the-box experience. "Everything is a plugin" turns Claude Code's baked-in assumptions into replaceable config — but freedom isn't free.

Two Philosophies

Harness : Infrastructure open-sourced. Kernel: Cordis handles only plugin load/unload/dependency management, no privileged core . Extension: everything is a plugin, replaced at config layer. Metaphor: shell workshop, assemble your own furniture.

Claude Code : Product perfected. Vertically integrated, deeply tuned for Claude model. Extension: Skills, MCP, Hooks at periphery; core closed. Metaphor: furnished apartment, move-in ready.

Three Wins for Harness

Forkless capability replacement : To run subprocesses outside the local machine, Claude Code needs wrappers or forks. Harness's capability seams are true alternatives: implement a provider, write to cordis.yml, and filesystem, subprocess, LSP all switch without touching core.

Auditability as architecture, not feature : Claude Code relies on internal context compression; you cannot precisely reconstruct what the model saw at a given step. Harness's append-only log plus "model-visible implies recorded" lets you exactly reconstruct any turn's model view, supporting fork/resume/search/replay. For teams with audit/compliance needs, this is the strongest argument.

Model-agnostic + private deployment : No vendor lock-in; fully local with local models. Claude Code binds to Claude series, no private deployment.

Strategic bonus: runtime self-extension ( cordis_* ) allows agents to modify their own runtime — the engineering foothold for self-evolution; Claude Code's architecture has no such slot.

Costs Harness Pays

Performance overhead unquantified : Revertible effects require maintaining inverse-effect linked lists; high-frequency ops (every tool call registers/unregisters effects) overhead marked as "future work" in the paper. Tax for reversibility; Claude Code's monolithic impl avoids it.

No out-of-the-box optimal defaults : Claude Code's diff preview, permission re-confirmation, context compression are carefully tuned defaults. Harness defaults are just a plugin combo; achieving parity requires selecting plugins, tuning config, even writing plugins.

Steep learning curve : Must understand effect/coeffect, fiber, ctx.effect(), inject, profile/bundle — the Cordis paradigm. Claude Code developers barely need these concepts.

Ecosystem fragmentation risk : "Too free" leads to uneven quality, compatibility reliant on community discipline. 288 plugins in 24 hours shows heat, but stability and security auditing are separate.

Maturity and debuggability : v0.1 explicitly warns of breaking changes, not production-ready. Dynamic composition makes root-cause harder (which plugin changed what at which moment).

Distributed scenarios uncovered : Paper's formalization targets single-process dynamic composition; cross-machine multi-agent composability across network boundaries remains open.

Comparison Matrix and Selection Guide

Dimension comparison :

Open source: Harness fully MIT (framework + ecosystem); Claude Code client open, model inference closed.

Architecture: Harness everything plugin, no privileged core; Claude Code closed core + open extension points.

Model: Harness fully decoupled, swappable (~40); Claude Code hard-bound to Claude series.

Replaceability: Harness model/tools/skills/session/sandbox/UI all swappable; Claude Code only peripheral.

Runtime: Harness hot-swap + self-extend ( cordis_*); Claude Code fixed.

Auditability: Harness append-only log, fork/replay; Claude Code internal context compression.

Private deploy: Harness yes; Claude Code no.

Out-of-box: Harness low (self-assemble); Claude Code high (polished product).

Maturity/performance: Harness overhead unquantified, v0.1 preview; Claude Code mature, enterprise-grade.

Scenario recommendations :

Production, stability, best CLI experience → Claude Code.

Self-built agent stack, private deploy, high audit/compliance → DeepSeek Harness.

Benchmarking different models in identical tool env → Harness (model-agnostic + minimalist mode).

Experimenting, learning agent substrate architecture, researching self-evolving agents → Harness (not for production).

Conclusion

Harness's value has two layers. As a product, it's still a shell: fast iteration, breaking changes, production caution. As infrastructure and strategy, it opens the question "how should agents be built" — open weights solved "who can run models"; open Harness tackles "who decides how models work." Claude Code proves how good a polished agent product can be today; DeepSeek bets that the agent era's substrate needs an open, auditable, self-evolving standard. The former wins now; the latter bets on the future.

For everyday developers, Harness's immediate appeal is a controllable, customizable, auditable, model-agnostic agent runtime — a Lego toolbox for plugging LLMs into local files and tools. Even if you return to Claude Code for daily work, spending an evening taking Harness apart is worthwhile.

References

GitHub: https://github.com/deepseek-ai/deepseek-harness (README.zh.md, docs/architecture.md)

Official site: https://www.deepseek.com/harness/

Cordis design paper: https://github.com/cordiverse/paper ( A Programming Paradigm for Spatiotemporal Composability , Peking University + DeepSeek-AI)

Paper analyses: Quantum Bit "DeepSeek's Self-Evolution Blueprint", Mushroom Research Blog, BestHub Cordis paper breakdown

Comparative benchmarks: Rohit Raj "DeepSeek Harness vs Claude Code vs Codex CLI", CSDN/53AI series

Media coverage: Global Times, Caixin, Cailian Press, 36Kr, Jiemian, Tencent Tech

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

AI AgentsPlugin Architectureopen-sourceAgent FrameworkClaude CodeCordisSpatiotemporal ComposabilityDeepSeek Harness
Tencent Architect
Written by

Tencent Architect

We share insights on storage, computing, networking and explore leading industry technologies together.

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.