clawpatch: AI Reviews AI Code via Local Agents — Pipeline with Full Audit Trail

clawpatch orchestrates local coding agents like Codex CLI and Claude Code as subprocesses to review, fix, and revalidate AI-generated code through a six-stage pipeline (init, map, review, fix, revalidate, open-pr) with semantic slicing, evidence-backed findings, read-only sandboxes, and explicit human approval — without ever calling model APIs directly.

Geek Labs
Geek Labs
Geek Labs
clawpatch: AI Reviews AI Code via Local Agents — Pipeline with Full Audit Trail

The Problem: Noisy AI Code Reviews

The author recounts using Codex CLI to review a legacy repository: it produced 47 "potential issues" spanning style, memory leaks, and missing comments. After ten minutes of inspection, 43 were false positives, duplicates, or lacked clear justification for why they were problems.

What clawpatch Solves

If coding agents (Codex, Claude Code, Cursor) are "interns writing drafts," clawpatch is the "senior engineer grading the intern." It does not write code; it reviews AI-written code, proposes concrete fixes, and optionally applies patches — all with a full evidence trail.

Unlike traditional linters that rely on rule matching, clawpatch slices the entire repository into "semantic feature slices" (feature slicing) and invokes the locally installed coding agent (Codex CLI, Claude Code, Cursor Agent, Grok Build, OpenCode, Pi) to review each slice. Every bug becomes a finding with an evidence chain, then the tool decides whether to patch and open a PR.

Core Design: No Direct Model API Calls

Most AI code-review tools call model APIs themselves — managing keys, billing, routing, retries. clawpatch reverses this: it never calls a model API. It only spawns the user's already-installed coding agent as a subprocess via the ACP protocol, letting the agent use its own sandbox, permissions, and model switching.

This principle is codified in VISION.md: "Pull requests and issues proposing direct model API providers are out of scope and will be closed." Consequences:

API keys, billing, model routing, retry policies — all handled by the coding agent, not clawpatch.

Users get the full capability of their chosen agent (sandbox, permissions, model switching), not a crippled wrapper.

clawpatch's provider adapters stay tiny (dozens to hundreds of lines each).

This avoids the "AI tool wrapper" trap where a project becomes an AI infrastructure company managing keys, billing, and stability.

Six-Stage State Machine

The workflow is a strict sequence:

init → map → review → next/show/report → fix → revalidate → open-pr

init (initialize) : Detects project type, writes .clawpatch/config.json and project.json; default mapper does not call a model.

map (slice) : Partitions the repo into semantic feature slices. Examples: Next.js by route, Go by package+binary, Rust by crate+binary. Supports Node/TS, Python, Ruby, PHP, Go, Rust, C/C++, Java, Kotlin, .NET, Swift. Generated directories and symlinks are skipped automatically.

review (review) : Runs the agent concurrently on several slices ( --limit 3 --jobs 3). Each bug is stored as a finding. Review runs in the agent's read-only sandbox (Codex: read-only sandbox; Claude Code: read-only tool allowlist; ACP via acpx --approve-reads).

next / show / report (inspect) : Findings are sorted by priority; each carries evidence — not "I think there's a bug" but "in file X line Y, function Z shows pattern W, and commit history shows this evolution."

fix (patch) : Triggered explicitly with --finding <id>. The agent applies a patch in a workspace-write sandbox. Default rejects dirty worktrees; no auto-commit. Every attempt plus validation result is stored under .clawpatch/patches/.

revalidate (re-check) : Re-runs the same evidence set to confirm the patch fixed the original finding without introducing regressions.

open-pr (open PR) : Separate command; not auto-run from fix. Commits, pushes a branch, opens a PR — does not merge.

All state lives in .clawpatch/; interrupted runs resume. Multiple processes coordinate via feature locks; stale locks auto-recover.

Complete Review Flow Example

cd your-project
clawpatch init          # detect project + write config
clawpatch map           # slice semantically
clawpatch status        # view slice status
clawpatch doctor        # check provider readiness (Codex CLI installed? authenticated?)
clawpatch review --limit 3 --jobs 3  # concurrent review of 3 slices
clawpatch report        # view finding report
clawpatch next          # see next finding to fix
clawpatch fix --finding <id>      # apply patch (evidence-backed, auditable)
clawpatch revalidate --finding <id> # re-check with same evidence
clawpatch open-pr --patch <id>    # explicit commit + push + open PR

CI can run a source-read-only loop:

clawpatch ci --since origin/main --output clawpatch-report.md

Safety Boundaries (More Important Than Features)

Safety is one of four core principles in VISION; docs/safety.md details constraints:

Review/revalidate default read-only — provider runs in read-only sandbox or read-only tool allowlist.

Fix must be explicit — requires --finding <id>; no "AI secretly changed your code."

Fix rejects dirty worktree by default — you must commit cleanly first.

Provider output must pass schema validation — malformed agent JSON is rejected.

No auto-commit, auto-PR, auto-merge, rollback snapshots, or global process locks — deliberately omitted.

Git safety remains the caller's responsibility: "Inspect git diff and run project tests before committing." This contrasts sharply with tools that chase full automation (auto-merge, auto-deploy); clawpatch instead asks: how to keep the agent inside the strictest boundaries doing the least possible actions.

Comparison with Similar Tools

Codex CLI / Claude Code direct review

Approach : Give entire repo to agent at once

Difference from clawpatch : No slicing, no finding storage, no revalidate; one-shot output, evidence chain broken

SonarQube / DeepSource / Codacy

Approach : Rule matching + historical trends

Difference from clawpatch : Rules are fixed; find known patterns; cannot catch AI-specific "looks right but semantically wrong" issues

Traditional PR bots (Sourcery, Codeball)

Approach : GitHub webhook → auto-run

Difference from clawpatch : Usually call model APIs themselves (manage keys); no slicing; review/fix not separated

Aider / Continue (AI coding tools)

Approach : Help you write code

Difference from clawpatch : Not responsible for review and patch; self-review is awkward

clawpatch occupies a unique niche: it is not an "AI programming tool" but an "AI programming tool reviewer + patcher."

Who It's For / Not For

Good fit:

Teams already using Codex CLI / Claude Code / Cursor Agent who want automated review + bug-fix pipeline.

Multi-language monorepos needing one tool for Node/Python/Go/Rust/Swift full stack.

CI pipelines needing "source-read-only continuous review reports" ( clawpatch ci --since origin/main).

Teams wary of "AI auto-fixing code" who want review and fix fully separated with complete evidence chains.

Poor fit:

No local AI coding agent installed (prerequisite).

Want "AI fixes all bugs then auto-merges" — clawpatch explicitly refuses auto-commit/merge.

Want to call model APIs directly without a coding agent — VISION closes such PRs.

Noteworthy Security Details

clawpatch doctor

runs before init — loads config then checks .clawpatch/ state. This allows injecting config via --config or CLAWPATCH_CONFIG before initialization, which is critical for CI integration.

Trusted config validation refuses to read provider.codexConfig from auto-discovered repo or state configs — preventing a malicious checkout from hijacking your Codex routing or credentials. Config only loads from explicit --config or environment variables.

Such "seemingly small, actually critical" security designs appear throughout. VISION summarizes: "expose the harness's security boundary honestly rather than claiming stronger isolation than the harness provides." This honesty about boundaries is the biggest differentiator from most AI tools.

Project Metadata

Repository: openclaw/clawpatch (TypeScript, MIT, 814+ stars, updated 2026-09). GitHub: github.com/openclaw/clawpatch. Note: openclaw here is a namesake collision — this project comes from the OpenClaw AI assistant tool team, unrelated to the OpenClaw platform mentioned in the article's source.

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

AI code reviewCI integrationCodex CLICoding Agentsclawpatchevidence-based findingsread-only sandboxsemantic slicing
Geek Labs
Written by

Geek Labs

Daily shares of interesting GitHub open-source projects. AI tools, automation gems, technical tutorials, open-source inspiration.

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.