DeepSeek Harness: Three Guardrails for AI Agents — Approval, Sandbox & Permission Presets
This article details Harness's three-layer guardrail system for AI agents: approval gates that default to deny, filesystem sandboxes with enforceable isolation levels, and permission presets that bundle sandbox and approval settings into switchable profiles, plus subagent and workflow orchestration capabilities.
1. First Layer: Approval — Should This Operation Proceed?
Approval answers a simple question: is this specific operation allowed to continue? In Harness, operations requiring approval enter an approval/request waterfall — a "show of hands" where the responder and the decision are determined by plugins.
Approval Results Are Closed
Approval has only four outcomes, and default is deny : allowed-once: Allow, but only for this single operation rejected: Deny cancelled: Request was withdrawn unavailable: No one could answer (default, also deny)
Note the last one: no response = deny . This is "fail closed" — without explicit authorization, the request cannot be treated as allowed. If an approval plugin is misconfigured or throws an exception, the effect is "all requests get no approval" — i.e., all denied.
Two Per-Session Strategies
ask : Normal approval flow; if someone answers, use that result; if no one answers, unavailable (deny). This is the default.
never : All approval-required requests are directly denied without even calling responders. Used for CI, unattended scenarios to strictly disable any operation needing approval.
Strategy changes are written to the session log, so replay can fully reconstruct the decision — no inconsistency where "it was allowed then but denied on replay."
What an Approval Request Looks Like
interface ApprovalRequest {
readonly agent: Agent // Which agent is requesting
readonly toolName: string // What tool is being requested
readonly callId?: CallId // Which tool invocation
readonly reason?: string // Why approval is needed (for humans)
readonly signal?: AbortSignal // Can be cancelled mid-way
}An approval request carries the agent, tool name, call ID, reason, and abort signal. Tool parameters are not part of ApprovalRequest; the call ID links the approval to the concrete tool call.
Who Responds?
Responders are approval/request waterfall listeners. Anyone can register:
Web UI : Pops up a dialog for the user to click Allow/Deny.
ACP (Agent Client Protocol) : Parent agent decides for child agent.
Custom approval plugins : Auto-approve based on rules (e.g., "only allow read, not write").
The first responder to return a result "takes the slot"; subsequent responders are not called — that's the waterfall behavior.
Audit Events
Every approval produces a pair of events: approval/asked and approval/decided, paired by the same ApprovalRequestId. These events are written only to logs, not to the model's conversation history — the model sees tool results and current permission state, not the approval process itself.
The core principle of this layer: every approval-required operation has an explicit result, an auditable trail, and defaults to deny.
2. Second Layer: Sandbox — The Filesystem Wall
Approval governs "whether to do it"; sandbox governs "what it can touch." Harness's sandbox focuses on one thing: filesystem effects . Network, process visibility, etc., are outside the sandbox's definition.
Three Modes
read-only: Read-only access. workspace-write: Can write within the workspace. danger-full-access: Bypasses filesystem sandbox protection entirely. The scary name is intentional — to make you pause when using it.
Enforcement Completeness: full vs partial
Sandbox has an "enforcement completeness" concept: full or partial. full: Backend fully implements all filesystem isolation promised by the mode. partial: Due to OS or kernel version limits, only partial isolation is achievable.
For example, newer Linux Landlock ABI achieves full; older kernels only partial. Windows is partial due to ACL boundary issues. This information is exposed — if your scenario requires complete filesystem isolation, seeing partial should trigger caution; you cannot treat it as full.
Per-Invocation Policy
Sandbox policy is not "set once at process start"; it is re-resolved on every capability invocation , traveling with the call. Why? Because a single process may simultaneously run components with different permissions:
Bash running in read-only mode.
A restricted child agent needing workspace-write for its own state directory.
Their policies don't interfere; each passes its own. The backend doesn't need to maintain state — it just executes per the given policy.
No Backend Available?
If no sandbox backend exists on the current system, ctx.sandbox.confine() throws SandboxUnavailableError. It does not silently degrade to no isolation — again, fail closed.
3. Third Layer: Permission Presets — One-Click Profile Switching
Approval and sandbox are two independent knobs, but users don't need to manage them separately. Harness bundles them into Permission Presets , like a phone's "Battery Saver / Performance Mode" — pick one and go.
Default presets provide two tiers: workspace-write and danger-full-access; finer control allows custom combos like read-only + ask.
workspace-write : sandbox workspace-write, approval ask — normal mode: can write workspace, operations require approval.
danger-full-access : sandbox danger-full-access, approval never — bypass filesystem sandbox, no approval needed.
Note: danger-full-access 's "full access" mainly targets the filesystem sandbox here; network, process visibility, etc., are not part of the Sandbox mode definitions.
You can add custom presets in config, e.g., read-only-review (read-only + approval) for code review mode.
How Presets Switch
The permission/preset event records only the user's intent — "which preset was chosen" — then separately invokes the sandbox and approval setters. Both knobs update independently; only the one that actually changed appends an event.
Why not replace sandbox/approval events with presets? Because presets are a UI convenience — the execution layer still reads each knob's result. Benefits:
Two presets may share the same knob combo (e.g., review-mode and safe-mode both = read-only + ask).
User picks preset A but manually changes approval to never — the effective state becomes custom.
Replay and persistence rely only on knob events, not preset names. custom is a derived state, not a real preset value — it means "current knob combo doesn't match any preset." UI can show it but cannot switch to it.
4. Subagents: One Agent Not Enough? Spawn Another
After the safety trio, we look at inter-agent collaboration. The most direct way is Subagent : parent agent spawns a child agent, delegates a concrete task, child runs and returns the result.
Unlike the bash tool (executes one command), a subagent launches a full agent loop — it thinks, calls tools, does multi-step reasoning.
Six Backends, One Interface
They solve "where and via what runtime does the subagent run" without changing the Subagent calling interface. spawn: In-process launch of a brand-new child agent fork: In-process launch, based on parent agent's completed history acp: Connect to external agent via Agent Client Protocol codex: Launch child agent via Codex app-server claude-code: Launch child agent via Claude Agent SDK dsh-sdk: Connect to another Harness agent via TypeScript SDK
Despite different backends, all register into ctx.subagents and are invoked via the same interface. The model sees a single start_subagent tool; backend choice is config-driven.
Capability Declaration
Each provider declares supported capabilities:
interface SubagentCapabilities {
readonly outputSchema: boolean // Supports structured output
readonly depthLimit: boolean // Supports depth limiting
readonly toolFilter: boolean // Supports tool filtering
readonly persona: boolean // Supports persona setting
}On launch, capabilities are validated — if the backend doesn't support a requested capability, it errors immediately, no silent degradation. Again, fail closed.
Depth Limit
Can a subagent spawn further subagents? Yes, but with a depth limit. You can set maxDepth, e.g., 2 layers, preventing infinite nesting. Depth is lineage-inherited: parent at layer N, its child at N+1.
Continuable Subagents
Subagents aren't necessarily "one-shot calls." Harness provides continuable subagents: state can be persisted, then you can send more messages, inspect status, or interrupt. You can list all subagents in the session, view their states, send messages, interrupt them — these operations are themselves tools the model can call.
5. Workflows: Let the Model Write Its Own Orchestration Script
Subagent = "spin up an agent for one task." But what if the task is complex — needs multiple agents, specific ordering, handling intermediate results? That's where Workflow comes in.
Workflow's essence: let the model write a script that can spawn subagents, await results, do logic; the engine executes the script.
// A workflow script roughly looks like this
const research = await agent("research", "Research competitor features");
const design = await agent("design", "Design solution based on research report");
const review = await agent("review", "Review design solution");
return {
research: research.output,
design: design.output,
review: review.output,
};Model writes script, engine executes; each agent() call truly spawns a child agent.
Common Confusion: Workflow's VM ≠ Filesystem Sandbox
Workflow uses Worker Thread + VM to run model-generated scripts, but this is not the same as the Sandbox discussed earlier. Sandbox constrains filesystem effects of concrete capability calls; Workflow's script execution environment is a separate runtime layer.
Therefore, do not equate "script runs in VM" with "script already has a full security sandbox." If running untrusted model-generated scripts in production, you must further verify the current Workflow Runtime, Worker isolation, and host permission boundaries — cannot judge safety solely by vm or worker_threads.
Script Runs in Worker Thread
Workflow scripts don't run on the main thread; they run in worker_threads — one Worker per workflow. Script executes in a VM context; Workflow Runtime exposes only a few built-in entry points: agent(label, prompt) — spawn a child agent log(message) — record log phase(title) — mark entering a phase (for progress display) args — input parameters
Script returns a JSON value on completion — that's the workflow result.
Why Not Let Model Call Subagent Tool Directly?
Because complex multi-step collaboration via "model says one step, execute one step" is slow and expensive — every step incurs an LLM inference round. Workflow's advantage: model writes the script once; thereafter the engine executes the orchestration without further LLM calls. Orchestration belongs to the engine; decision-making belongs to the model; each does its job.
6. More Platform Capabilities at a Glance
Goal
Give an agent a goal and a max-turn limit. After each turn, the agent checks if the goal is achieved; if so, it stops; hitting the limit also stops. Goals have their own lifecycle: active / paused / blocked / complete. Blocked includes a reason code and description, routable to a human.
Jobs
Jobs are Harness's unified runtime mechanism for managing long-running operations. Some time-consuming capabilities use Jobs for lifecycle management, gaining unified status query, cancellation, and output reading.
Skills
Like Codex skills — Harness has a skill system. Skills are instructional text (not events), loadable from local dir, project dir, user dir, bundled, etc. Model uses the skill tool to load a skill, then follows its instructions. Skill registry is layered (host layer + scope layer), same structure as tool registry — nearest same-named skill wins.
Summary
This article boils Harness's permission system down to three keywords:
Approval : Controls "should this operation proceed" via approval/request for one-time authorization with fail-closed semantics.
Sandbox : Controls "how far can the operation reach" via read-only, workspace-write, danger-full-access describing filesystem boundaries.
Permission Presets : Bundles the two independent knobs (sandbox + approval) so users can switch permission modes with a single profile.
Above that are Agent collaboration capabilities:
Subagent : Delegate a task to another agent.
Workflow : Model generates orchestration script; engine executes multi-agent collaboration.
Goal / Jobs / Skills : Respectively solve goal management, background execution, and capability injection.
So Harness doesn't just give agents "hands"; it also draws boundaries around those capabilities.
The real question isn't how many tools an agent can call, but:
As agents grow more capable, who decides what they can do, how far they can go, and when they must stop and ask a human?
That may be one of the biggest differences between an Agent Runtime and a plain LLM SDK.
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
AI Code to Success
Focused on hardcore practical AI technologies (OpenClaw, ClaudeCode, LLMs, etc.) and HarmonyOS development. No hype—just real-world tips, pitfall chronicles, and productivity tools. Follow to transform workflows with code.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
