Beyond the Model: Dissecting the Harness Layer of 11 Coding Agents
This article analyzes the harness layer of 11 coding agents—including Claude Code and Codex—based on an 83-page arXiv paper, revealing how runtime systems manage tools, context, permissions, and verification beyond the model itself.
Source-Code Study of Eleven Coding Agents
An 83-page arXiv paper titled Harness Engineering: Anatomy, Architecture, and Evolution of Coding Agents (Barbaste et al., 2026) examines the source code of 11 coding agents: Claude Code, Codex CLI, Gemini CLI, Mistral Vibe, OpenHands, Aider, Mini-SWE-Agent, Hermes, Pi, OpenCode, and OpenClaw, using Omnigent as a meta-harness reference. The analysis covers a March 2026 snapshot of Claude Code’s leaked source and release logs—not a full audit of the current closed-source version. The paper extracts 13 cross-cutting observations, 29 reusable patterns, and 18 design recommendations, plus a ~90-line minimal harness skeleton.
Harness as the Runtime Control Plane
The harness sits between the model and the engineering environment. It prepares context, exposes tools, executes actions, enforces permissions, persists state, compresses history, coordinates subtasks, and decides whether to continue based on verification signals. The paper decomposes the harness into seven responsibility boundaries:
Agent Loop : how the next step advances and when to stop.
Model Integration : handling different model protocols, streaming, caching, and errors.
Tools & Actions : what can be read, edited, executed, and how results are returned.
Memory & Context : what the model sees this turn, handling history overflow, and cross-session persistence.
Safety & Permissions : whether an action is allowed and what resources the execution environment can touch.
Orchestration : how sub-agents divide work and how results are aggregated.
Extensions : integrating Skills, MCP, Hooks, plugins, and client protocols.
The runtime loop runs: Goal (user intent) → Context (current view) → Decide (model proposes next step) → Act (execute) → Observe (return result) → Verify (check completion) → State (persist events) → Next (continue, stop, retry, or hand off).
The Loop Is Short; the State Is Hard
The basic loop—assemble input → call model → execute tool → return observation → repeat—is trivial. The differences lie in how each system handles state, concurrency, cancellation, recovery, and verification.
Mini-SWE-Agent : linear loop, single shell tool, minimal guards—good for research and seeing the minimal skeleton.
OpenHands : persistent event log organized around sessions, branches, and replay; tool calls acquire resource locks (reads don’t block reads, writes to same file queue).
Claude Code (per the 2026 snapshot): batches tools by concurrency suitability—reads and searches run aggressively in parallel; writes and command execution are more conservative.
Codex CLI : asynchronous state machine separating session, model streaming events, tool execution, and external event reads—practical for CLI, TUI, SDK, and server entry points.
The snapshot also shows event-layering: remote sessions distinguish normal SDK messages, permission requests, permission cancellations, user input, and interrupt signals; tool progress and compression boundaries are emitted as events. This supports the paper’s “event-driven control plane” claim.
Tools: Contract Clarity Over Quantity
Every tool definition consumes context tokens (name, parameters, description, constraints). More tools increase selection burden and mis-selection risk. Anthropic’s Building Effective Agents advises fewer, clearer, composable tools.
Claude Code’s context advisor checks token usage of Bash, Read, Grep, WebFetch results and suggests narrowing with head, tail, grep, offset, limit.
Edit tools illustrate the contract spectrum:
Exact-match : require old-text and new-text; old-text must match uniquely. Fail on no match or multiple matches. Hard contract, but failures are precise—model must re-read if it misremembered.
Fuzzy-match : tolerate whitespace/indentation drift. Higher success rate in some cases, but ambiguity when two segments look alike—which one was changed?
Systems choose differently: Aider offers multiple edit formats with correction flows; Codex uses patch format; Hermes and OpenCode emphasize match tolerance; Pi handles Unicode/whitespace diffs and checks overlapping edits; Mistral Vibe moved from lenient SEARCH/REPLACE to strict unique old-text matching in its audit window.
The harness must distinguish at least three failure-mode pairs so the model can recover:
No search results vs. search output truncated .
Command exited with failure vs. command still running .
Edit found no unique location vs. edit applied but tests fail .
An inspected MCP tool implementation returns “not found”, “multiple matches”, “try whitespace normalization” for edits; command tool distinguishes timeout, exit code, long-running session, and process-not-found. These details determine whether the model re-reads, retries, waits, or escalates to a human.
Context Compression Is State Preservation, Not Summarization
Long tasks inevitably hit context limits. Compression must answer: what minimum state does the next turn need? If compression leaves only “implementing CSV export”, critical constraints like “reuse current filter” and “don’t break pagination” are lost.
Aider : summarizes older history via the model, keeps recent raw text (recent error stacks are more useful than “compilation failed”).
OpenHands : makes compression a Condenser ; compression requests and results are themselves recorded in the event log, enabling post-hoc debugging of what was dropped.
Gemini CLI : generates state snapshots and retains recent history.
OpenCode : merges summaries into structured fields—goal, key details, current state, next step, relevant files.
Hermes : links parent/child sessions to preserve pre/post-compression relationships.
Pi : append-only JSONL tree enables natural session forking and review.
Two distinct concerns: context serves next-turn reasoning ; event log serves recovery, replay, and audit . Just because a log segment isn’t in the current context doesn’t mean it’s useless.
Long-term memory should store reusable cues (test commands, directory conventions, stable architectural constraints)—not transient guesses or fixed bugs. Mature harnesses give memory provenance, permissions, and correction paths; otherwise stale premises accumulate.
No Vector Retrieval in the Sample—A Reminder
The paper found no runtime vector-embedding code retrieval in the 11 systems. This is a sample observation, not a universal claim. But it warns against defaulting to RAG for code tasks. Code has strong structural signals: paths, filenames, symbols, types, imports, test names, stack traces, call graphs. Tools like ripgrep, glob, tree-sitter, language servers, and RepoMap often locate the right spot.
Aider’s RepoMap uses tree-sitter to extract symbols, ranks them by task relevance, and fits a repository map into the token budget. This is structural navigation, not full-text search. The safer order: exploit deterministic signals (paths, symbols, text, diagnostics) first; only if real tasks repeatedly fail because users provide only business descriptions without keywords should semantic indexing be evaluated.
Three retrieval types must be separated:
Source-code retrieval : find code in the current workspace.
Conversation-history retrieval : find what happened before.
Long-term memory retrieval : find reusable rules and preferences for future tasks.
OpenClaw’s hybrid memory retrieval doesn’t prove “coding agents should use vector search for code”. Different problems demand different strategies.
Permissions and Sandboxes Are Not the Same
User consent (“I agree to run tests”) authorizes the action, but not the side effects: what files the test process reads, what network it accesses, what subprocesses it spawns.
The paper compares sandboxing and permission models concretely: Codex invests heavily in cross-platform sandboxes; Gemini CLI has cross-platform isolation; Claude Code offers optional OS-level sandbox; OpenHands uses workspaces and multiple execution backends; others rely on command permissions, checkpoints, or content threat detection. System size ≠ sandbox strength; feature count ≠ clear side-effect boundaries.
Four distinct questions:
Authorization : is this action allowed?
Isolation : if allowed, what resources can it affect at most?
Recording : are side effects evidenced?
Recovery : on retry, will side-effecting actions re-execute blindly?
“Stop” must resolve all four: did the model request stop? did the tool stop? did child processes stop? did parent stop while children still consume tokens, edit files, run commands? “UI shows stopped” ≠ “system actually stopped”.
Multi-Agent: Boundaries First, Count Second
Multiple agents ≠ “more hands = faster”. Long tasks don’t always benefit from parallel agents; many modules don’t always allow parallel edits. Example: CSV export needs frontend and backend changes, but the API parameters aren’t confirmed. Two agents writing simultaneously will likely mismatch. Better: one investigation task to confirm the interface, then separate implementation, then unified integration test.
Orchestration patterns vary: Claude Code and Codex have coordinator modes; Mistral Vibe uses sequential delegation; Gemini CLI has an agent registry and remote session interface; Hermes intersects sub-task permissions; Pi spawns independent processes via extensions; OpenHands and OpenCode use independent sessions or sub-sessions. All are called “multi-agent” but serve different goals: parallel exploration, context isolation, remote service reuse, outsourcing a local investigation.
“Supports sub-agents” ≠ “faster”. Sequential delegation may have no parallel gain but keeps context clean. Parallel workers may save wait time but introduce conflicts, merge issues, and budget problems.
What matters is sub-task boundary clarity. “Look at the backend” is vague; “check which filter parameters the export endpoint supports, return file locations, parameter definitions, unsupported fields, do not modify code” is precise. Sub-agent tools can be scoped by phase: investigation gets read/search; implementation gets edit/execute. The coordinator then receives evidence, not a raw environment.
Skills, MCP, ACP: Don’t Mix Them
Three extension interfaces often conflated:
Skills : process assets—steps, rules, domain knowledge, optional scripts for a class of tasks. Adoption in the sample: 9/11 systems, higher than MCP (8/11). Teams often teach agents “how we do things here” before connecting external systems.
MCP (Model Context Protocol) : standardizes exposing tools and data services (databases, internal systems, filesystems, browsers) to the host.
ACP (Agent Client Protocol) : client/host interface. In the sample, 6 systems implement ACP, and it’s evolving beyond editor↔agent to harness hosting—one system can host another harness via an adapter.
This shift from tool-layer to platform-layer is why the paper discusses harness as a platform runtime. A CLI serving only itself has simple boundaries; once it offers SDK, HTTP, session APIs, plugin/skill marketplaces, and external agent hosting, the harness becomes an embeddable, extensible, governable platform. The platform isn’t the chat UI—it’s how each turn is assembled, executed, observed, and taken over.
Meta-harness, multi-layer feedback loops, and context engineering point to the same trend: stop staring at prompts, start asking “who maintains the workbench for each turn?”.
Building a Harness: First Five Priorities
The 18 recommendations need not be adopted all at once. A code-task harness can start with five foundations:
Event logging : model inputs, tool parameters, tool results, file changes, test results, termination reasons—all queryable. Without it, failure debugging falls back to guessing from chat history.
Clear tool contracts : edit tools define locating, rejecting ambiguity, reporting failure; command tools distinguish running, timeout, exit failure, output truncation; search tools distinguish no-results from incomplete-results.
Context compression : summaries must retain goal, constraints, changed files, verification results, current blocker, next step. Compression must not delete sole evidence; full event log stored separately.
Permission and execution boundaries : user consent, tool permissions, file scope, network scope, subprocesses, cancellation semantics—designed separately. Side-effecting operations must not blindly re-run on recovery.
Verification loop : code written ≠ task done. At minimum, wire back tests, lint, type checks, diffs, acceptance criteria, and human confirmation. Without verification, the agent just produces “looks done” faster.
This echoes Martin Fowler’s outer-harness discussion: pre-action constraints (AGENTS.md, Skills, architecture rules, project conventions) and post-action checks (static analysis, tests, structural rules, review agents, human review). Internal runtime and external team processes are not isolated; both affect whether the agent can self-correct and the team can avoid wasteful re-reviews.
Only after these are solid should Skills, MCP, multi-agent, and remote hosting be added. Extensions amplify existing capabilities—and existing chaos. A single agent that hasn’t nailed state, permissions, and verification will only make problems harder to trace when split. Conversely, once real failures are recorded, explained, and replayable, extension points have a foundation.
What the Harness Must Ultimately Bear
The paper’s value isn’t “copy this system” but decomposing the coding agent’s runtime responsibilities: model integration must handle streaming, caching, errors, model differences; tools must define recoverable failure states; context compression must preserve task constraints and working set.
Permissions cannot stop at user confirmation. Execution isolation, side-effect recording, cancellation semantics, and recovery strategies must be designed in. Multi-agent and extension interfaces follow the same logic: clarify boundaries first, then decide on Skills, MCP, ACP, or sub-agents.
These questions are closer to the engineering floor than “which framework”, “how many tools”, “how many agents”.
Two final observations to weigh calmly: no general-purpose agent framework appeared in the runtime sample, and no vector embedding code retrieval. This isn’t declaring frameworks or vector search useless—it’s reminding us that production coding agents derive critical capabilities from deterministic systems engineering: loops, events, tool contracts, context, permissions, verification.
These pieces aren’t flashy, but they determine whether an agent graduates from “can talk, can write” to “can take over, can recover, can deliver”. The stronger the model, the less these engineering details should be treated as peripheral glue. Once the model acts, every undefined boundary lands in files, processes, permissions, costs, and delivery outcomes.
The stronger the model, the more the harness must bear load—not just wrap it.
References
Paul Barbaste, Tristan Darrigol, Germain Vu, Tom Wiltberger, Harness Engineering: Anatomy, Architecture, and Evolution of Coding Agents -- A Source-Code Study of Eleven Systems (https://arxiv.org/abs/2609.00006)
Anthropic, Building effective agents (https://www.anthropic.com/engineering/building-effective-agents)
Anthropic, Effective context engineering for AI agents (https://www.anthropic.com/engineering/effective-context-engineering-for-ai-agents)
Anthropic, How we built our multi-agent research system (https://www.anthropic.com/engineering/multi-agent-research-system)
Model Context Protocol, Specification 2025-06-18 (https://modelcontextprotocol.io/specification/2025-06-18)
Simon Willison, Context engineering (https://simonwillison.net/2025/Jun/27/context-engineering/)
Martin Fowler, Harness Engineering (https://martinfowler.com/articles/exploring-gen-ai/harness-engineering.html)
Andrej Karpathy, Context engineering (https://x.com/karpathy/status/1937902205765607626)
Andrew Ng, Loop engineering & multi-layer feedback loops (https://x.com/AndrewYNg/status/2071988145667928442)
Harrison Chase, Meta-Harness & harness layer learning (https://x.com/hwchase17/status/2040471961206214864)
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
Architect
Professional architect sharing high‑quality architecture insights. Topics include high‑availability, high‑performance, high‑stability architectures, big data, machine learning, Java, system and distributed architecture, AI, and practical large‑scale architecture case studies. Open to ideas‑driven architects who enjoy sharing and learning.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
