Agentic Harness Workflow: Engineering AI Coding into a Fixed Pipeline with Specialized Sub-Agents

The author presents a self-built framework that decomposes AI-assisted development into a fixed 12-stage pipeline — from requirement clarification to archival — each executed by a dedicated sub-agent orchestrated by a central Manager, with state persisted to disk, human approval gates at critical steps, and git worktree isolation for parallel requirements.

Baidu Geek Talk
Baidu Geek Talk
Baidu Geek Talk
Agentic Harness Workflow: Engineering AI Coding into a Fixed Pipeline with Specialized Sub-Agents

The article introduces Agentic Harness Workflow , a framework designed to turn AI coding from a hit-or-miss activity into a repeatable engineering process. The author identifies four core problems with feeding an entire requirement to a single long-context agent: context bloat and drift, skipped engineering steps (clarification, design review, test design), non-recoverable process after interruption, and lack of human approval gates.

Why a Harness Is Needed

Instead of building a smarter agent, the solution is to impose an engineering framework: a fixed pipeline where each stage is handled by a single-purpose Sub-Agent, coordinated by a Manager that schedules, manages state, and translates between structured sub-agent protocols and natural-language user interaction.

Three Application Topologies Compared

The author contrasts three structural approaches conceptually, positioning the Harness as Workflow-centric, phase-level Multi-Agent collaboration with Human-in-the-loop . (Note: "Multi-Agent" here refers to the Sub-Agent capability built into tools like Claude Code and Codex, not fully autonomous agents.)

Overall Architecture

3.1 Complete Lifecycle

A requirement advances through a fixed sequence:

init → specify → plan → tasks → test-plan → plan-review → implement → code-quality → e2e-run → human-acceptance → commit-push → archive

The segment implement → code-quality → e2e-run runs automatically; most other stages pause for user confirmation. human-acceptance and archive have no independent agents — the Manager handles presentation, acceptance, and wrap-up.

3.2 CLAUDE.md: Framework Entry Point & Project Knowledge Base

CLAUDE.md

serves two roles:

Framework entry — auto-loaded at session start; it only declares the phase sequence, points to three rule files ( state-graph.json, protocol.md, state-schema.md), and mandates adherence to .harness/rules/manager.md.

Project knowledge base — indexes repo list, service startup, test DB connections, ticketing system, inter-repo dependencies. Only the Manager reads CLAUDE.md; it extracts relevant slices and passes them via named input fields (e.g., environment_knowledge to e2e-run, icafe_space to commit-push, knowledge_refs to plan). Benefits: (1) Sub-Agents don't swallow the whole knowledge base, (2) Sub-Agents stay generic — swap CLAUDE.md per project without touching agents/ or rules/, (3) knowledge changes in one place.

A hard-won principle: don't assume, ask if missing — repo configs vary (e.g., conf/servicer/*.toml), so CLAUDE.md acts as a knowledge-layer plugin while rules/ (skeleton) and agents/ (phases) remain universal.

3.3 Scheduler: Main-Session Manager

The Manager performs three functions:

Scheduling — decides which phase's Sub-Agent runs next, forward or rollback.

State management — sole writer of workflow-state; moves requirement directories between active/, process/, archived/.

Conversation output — Sub-Agents speak structured protocol; user speaks natural language; Manager translates both ways.

Iron rule: Manager only reads the Sub-Agent's structured protocol response (verdict: pass / reject / need-input) and the state file — never the full artifacts (spec, plan, code, reports) or business code. This keeps the main session light and stable across dozens of phases. If asked "why that verdict?", Manager re-delegates to the originating Sub-Agent for explanation.

Scheduling isn't purely heuristic; a state graph in state-graph.json defines legal transitions (each later phase can roll back to any earlier completed phase). Manager combines Sub-Agent suggestions with user input to decide.

3.4 Executors: Phase Sub-Agents

Each Sub-Agent is defined by its own Agent.md. The article includes a table (image) listing every phase, its purpose, and its artifact. Every phase communicates via a structured output protocol ; Manager presents the Sub-Agent's human-readable summary without reading the artifact itself.

A deliberate design separates "referee" from "player": design and review are split into plan and plan-review (similarly for test), inspired by adversarial thinking — reviewers challenge assumptions, not just verify happy paths.

3.5 State Machine: workflow-state

Process state lives in workflow-state.json per requirement, not in chat memory. Key fields: current_state: current phase pending: what we're waiting for (forward confirm, clarification, rollback confirm, human acceptance, push confirm, CR review) states: phase history, status, artifact paths review_findings: gate-detected issues and target rollback phase

Loop counters, involved repos, acceptance services, etc.

This enables Handoff : if you switch from Claude Code to Codex, or a session compresses/clears memory, the new session reads the state file and continues exactly where it left off.

3.6 Protocol: Manager ↔ Sub-Agent Communication

A single bidirectional protocol object:

{
  "state": "target phase (Manager fills)",
  "user_message": "user's raw words this turn (Manager fills)",
  "input": { "phase-specific context (Manager fills)": "..." },
  "status": "ok | needs_input | fail (Sub-Agent fills)",
  "output": { "structured result (Sub-Agent fills)": "..." }
}

Request side ( input ): per-phase keys predefined in manager.md scheduling rules — this keeps Sub-Agents generic. Three optional keys appear only in specific situations (images show rollback_reason, skip_gates, retry_count).

Response side ( status + output ): three statuses — ok (phase done, await user confirm), needs_input (human intervention needed), fail (phase failed, output.next_state suggests rollback target, e.g., e2e-run fail → implement). output carries four common fields plus phase-specific keys. Response is wrapped in a control-result block for Manager parsing.

Note: protocol.md , phase templates, and business code are read by Sub-Agents themselves, not passed by Manager.

3.7 Loop Engineering: Implement → Code-Quality → E2E-Run

Inspired by Codex's /goal and Loop Engineering, the framework auto-loops

implement → code-quality → e2e-run → implement (targeted fix)

but with a hard iteration cap . When the cap is hit, the loop stops and presents three choices to the user: continue fixing, force-pass, or ask the Sub-Agent to explain.

3.8 Phase Rollback: Root-Cause & Gate-Driven

All rollbacks follow one mechanism, triggered by:

User overturns upstream (e.g., realizes missing requirement at plan).

Downstream Sub-Agent reports upstream gap (e.g., test-plan finds undefined API routes in plan).

Gate rejects ( plan-review points to earliest faulty phase).

Auto-loop test failure ( code-quality / e2e-run → back to implement).

For human acceptance or CR review failures, Manager consults a fixed mapping table (image) to determine rollback target. Target must be earlier in

state-graph.json
order

and must have been genuinely completed in this requirement's history — preventing chaotic jumps.

On validated rollback, Manager performs an atomic state transition : current phase → terminal, target phase "completed" → "in-progress", current_state updated, and rollback_reason injected into target Sub-Agent's input so it knows what to fix. While awaiting user confirm, the current phase stays "in-progress"; terminal state is written only on actual transition — ensuring cold-start recovery always sees a correct state (at most one phase "in-progress", matching current_state).

3.9 Human-in-the-Loop: Explicit Confirmation Points

The pending field encodes every wait type with "what we're waiting for" and "next phase on confirm" (image table).

From Worktree to Acceptance: Isolating Concurrent Requirements

4.1 Workspace Directory Structure

<workdir>/
├── CLAUDE.md
├── .mcp.json
├── .harness/
│   ├── rules/ (manager.md, protocol.md, state-schema.md, state-graph.json, templates...)
│   ├── active/<idx>-<slug>/   ← at most one active
│   ├── process/<idx>-<slug>/
│   ├── archived/<idx>-<slug>/
│   └── worktrees/<idx>-<slug>/
│       ├── <repoA>/
│       └── <repoB>/
├── .claude/
│   ├── agents/ (init-agent.md, specify-agent.md, ... commit-push-agent.md)
│   └── skills/ (huibo-test-mysql, huibo-test-redis, icafe-card-assistant, icode, get-ugate-token)
├── <biz-repoA>/
└── <biz-repoB>/

Single-requirement directory ( .harness/active/<idx>-<slug>/) contains:

workflow-state.json
spec.md, spec_user_message.md
plan.md, plan_user_message.md
tasks.md
test-plan.md
reports/ (plan-review.md, code-quality.md, e2e-run.md, screenshots/)

4.2 Requirement Isolation

Uses git worktree. Before Implement, Manager reads the repo set from tasks output, creates a same-named feature branch and independent worktree per business repo. Implement, code-quality, and e2e reuse the same worktree; fix cycles don't rebuild; cleanup happens at archive. Benefits: (1) main workspace stays clean, (2) code review sees one consistent diff, not re-derived snapshots, (3) interrupt/rollback resumes with already-checked tasks intact.

4.3 Commit Dependency Handling

commit-push

splits into prepare (organize changes, create iCafe ticket, generate confirmation table) and execute (commit, push, create CR on user confirm). CR review stays human; Harness only proceeds or archives based on outcome. Archive stops acceptance services, moves requirement dir, deletes worktree and merged branches.

For multi-repo dependencies (common shared lib), dependency graph lives in CLAUDE.md. Manager feeds it to commit-push, which submits dependent-repo CRs first, waits for merge (manual +2 today, could automate via Skill), then dispatches implement with skip_gates to update dependents (e.g., go get xxx@latest) — skipping the full test loop for trivial dependency bumps.

DeepSeek Harness: "Everything Is a Plugin"

DeepSeek Harness (Aug 2026, MIT, CLI dsh) champions "everything is a plugin" — model, tools, skills, session, sandbox, storage, loop, scheduler, UI, even the agent loop itself are swappable plugins. Philosophy: "start from a working whole Agent and replace parts, not assemble from parts." Append-only session log enables resume/fork/replay.

The author argues this Workflow should also be a framework: minimal immutable core (Manager), everything outside replaceable file-plugins :

Each phase = plugin — one Markdown Sub-Agent file; add/remove/reorder by editing files and the flow table.

Skills = plugins — DB access, cache, ticketing, CR, Playwright — loaded by name.

Rules & knowledge = config — protocol, order, state schema, templates, project facts — all external files; change flow by editing flow table; new project = new knowledge base, core logic untouched.

Progress = recoverable — state-on-disk mirrors DeepSeek's append-only log for resumability.

Difference: DeepSeek Harness is a general Agent runtime (even loop swappable); this is a coding-lifecycle-specific Harness with a fixed pipeline , but inside that pipeline phases, skills, rules, knowledge are all pluggable — Manager minimal, phases & skills swappable .

Real Smoke Test

End-to-end run on a real requirement covering: normal flow, multi-round clarification, mid-flight requirement change + rollback, gate rejection + multi-level re-review, auto-loop fix cycles, cold-start resume, CR creation, archive. Includes a timeline screenshot and a detailed timing breakdown table (images).

Unsolved Problems & Next Steps

Sub-Agent reloads context every round — multi-round phases (specify, plan) re-read default context each turn. Idea: reuse Sub-Agent session ID (Claude Code does this for follow-up questions) to preserve context within a phase.

Plan phase has no search boundary, times out — observed 1-hour runs. Need knowledge-base integration with progressive indexing: limit repos, file count, call-depth, timeout, and force clarification when scope explodes.

Artifacts and state not atomic — e.g., test-plan.md exists and marked Ready but Manager can't verify completeness on re-entry. Must make "write artifact + update state" an atomic, verifiable step.

Manager oversteps answering user questions — during review, Manager sometimes reads code to explain instead of delegating back to the Sub-Agent that made the call. (Rule since tightened; issue no longer reproduces.)

Frontend design lacks guardrails — AI over-engineers UI; need design-system Skills or designer involvement.

No cloud-shared evolutionary memory — local memory lost on machine change. Want cloud memory so each run's bugs/fixes accumulate into "how to prevent next time" across machines/sessions.

Workspace structure not team-friendly — currently all in one git repo but not used that way. Explore git submodule for cloud storage; local workspace root shouldn't be a .git dir.

No production-grade observability — only phase artifacts and state snapshots exist. Need logs, traces, metrics, alerts for "did this phase execute correctly, read expected files, call expected tools?" — via custom Hooks or Coding Agent internal Hooks.

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

state managementAI codingsoftware engineeringgit worktreehuman-in-the-loopagent workflowsub-agentspipeline engineering
Baidu Geek Talk
Written by

Baidu Geek Talk

Follow us to discover more Baidu tech insights.

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.