Why ForgeX Stopped an AI Delivery Mid-Process — And What It Proves
The article details ForgeX's integration of a real project (Bracelet) which halted at 'action_required' due to incomplete Knowledge/Skill/MCP, demonstrating the platform's safe-stop design; separately, an end-to-end verification in PostgreSQL with a simulated requirement proved the full pipeline works when conditions are met, emphasizing that a releasable candidate is not a production release.
Chapter 1: Real Requirement Enters Platform — First Step Isn't Coding
1. A Seemingly Simple Requirement
The project integrated was Bracelet. The requirement was a style and experience adjustment around existing pages — a task that appears suitable for a Coding Agent to scan the repo, modify code, and run tests. The naive fastest path would be: enter directory → let Agent scan code → directly modify → output result. However, an enterprise delivery platform must first answer: which project, repo, branch? Is the workspace clean? Which directories are committed vs. untracked developer work? Are startup, test, build, and acceptance methods defined? What Knowledge, Skill, and MCP capabilities does the Agent need? Are those capabilities authorized for this task only? If conditions are incomplete, should the platform guess or stop and demand completion? The more capable the Agent, the greater the risk if these questions go unanswered.
2. Platform First 'Freezes the Scene'
ForgeX did not write code first; it inventoried the environment. On a standard server it identified 9 candidate source repos. Only source repos were imported — not runtime directories, build caches, database files, or credential-bearing env directories. This conservative act establishes a key boundary: project onboarding targets versionable source code, not every file that happens to exist on a server.
Project onboarding targets versionable source code, not every file that happens to exist on a server.
Further inspection of Bracelet revealed two critical facts: the current branch was refactor/miniapp, and the workspace contained untracked directories backend/, miniapp/, docs/. These untracked files may be uncommitted developer work or temporary artifacts; without evidence the platform cannot judge them, nor can it run destructive Git operations to 'clean up'. The output of this step is a project snapshot:
Repository Snapshot
├── repository identity
├── current branch
├── base revision
├── tracked changes
├── untracked paths
├── delivery profile
└── prohibited mutation scopeThe snapshot ensures every subsequent Agent starts from the same factual baseline. Without it, the same instruction 'modify Bracelet page' could map to different source, branches, or directories on different execution nodes.
3. standard-delivery@1 Is a Delivery Contract, Not a Prompt
After project import, ForgeX applied the standard-delivery@1 initialization rules. The delivery profile is not a long prompt for the model but a platform-checkable delivery contract. It requires three asset categories before a project enters Agent execution:
Knowledge : lets the Agent know project facts — architecture, directory responsibilities, startup methods, coding standards, domain terminology, known risks, acceptance criteria. Solves: Does the Agent understand this project?
Skill : lets the Agent work per team-approved methods — creating isolated workspaces, running frontend tests, generating DB migrations, building artifacts, collecting evidence. Solves: What engineering process should the Agent follow?
MCP : lets the Agent connect to external systems within controlled boundaries — reading requirements, querying the code platform, uploading artifacts, writing back status. Solves: Which real capabilities is this Agent allowed to call for this task?
For Bracelet, these three checks did not all pass. The final state was not ready but action_required.
Platform already knows what's missing, but current evidence insufficient to authorize execution.
Consequently, ForgeX did not copy local MCP configs to the server, did not copy any secrets, and did not enter the target source directory to modify code just to manufacture a 'success'. The source project, target directory, and other projects on the same machine remained untouched.
4. Why 'Stopping' Is the Most Important Result
Traditional automation treats stop as failure. Agent systems often make a different mistake: when information is missing, they fill gaps from similar projects, directory names, or general experience and continue acting. This may work for example code but is dangerous in real delivery. The missing piece might be authorization, a business decision, or a boundary only the project owner can confirm — not a default value. Therefore, an enterprise platform must treat 'cannot continue' as a first-class state: draft: requirement not yet stable → allow clarification, no execution assigned. awaiting_confirmation: candidate scope formed, waiting human confirmation → freeze candidate, wait approval. action_required: project conditions, permissions, or evidence incomplete → list gaps, forbid implicit bypass. ready: baseline and capabilities ready → allow scheduling. blocked: execution hit non-auto-resolvable block → preserve scene, return to human decision.
Platform's value is not letting all requirements proceed, but letting only those meeting conditions proceed.
Bracelet's real onboarding produced no candidate commit, yet it validated project-level boundaries, dirty workspace detection, initialization checks, and safe-stop mechanisms. If a company's AI platform can only showcase successes but cannot clearly show 'why it didn't execute', it remains far from real R&D governance.
Chapter 2: From Requirement Confirmation to Human Acceptance — How Multiple Execution Roles Relay
1. 'Multiple Agents' Are First Independent Responsibility Roles
In this context, Agent does not merely mean 'multiple LLMs chatting'. More accurately, they are multiple execution roles with independent identities, permissions, and responsibility boundaries:
Product Owner : clarifies requirement, confirms version, preview acceptance; must not forge Runner verification results.
Control Plane : solidifies state, schedules tasks, verifies tokens, organizes evidence; must not secretly modify source code.
Worker / Coding Agent : implements requirement in specified repo and isolated workspace; must not self-certify independent verification.
Runner : runs verification on specified commit, generates Preview and evidence; must not modify candidate code.
Release Owner : decides merge, release, canary, rollback; must not replace version-evidence binding with verbal judgment.
Worker can be Codex or another automated executor; Runner need not be an LLM — its key attributes are independence, determinism, and auditability. The real value of 'multiple Agents' lies in mutual constraints: implementer cannot stamp own verification; verifier cannot alter verified object; platform cannot skip evidence because a node says 'done'; final business acceptance remains with the decision-maker.
2. A Complete Chain Binds Six Objects
End-to-end verification starts with requirement creation. Product Owner adds scope and acceptance criteria; system forms requirement version; confirmation allows execution. The chain then proceeds:
Requirement version confirmed
↓
Transactional scheduling event registered
↓
Worker picks up leased execution task
↓
Isolated workspace produces candidate commit
↓
Runner verifies same repo, same commit
↓
Generates signed evidence and content-addressed Preview
↓
Product Owner previews and human acceptanceSix objects must be tightly bound: requirement identity, confirmed requirement version, project & repo identity, base & candidate commits, Runner test evidence, user-opened Preview artifact. If any binding breaks, 'misaligned completion' occurs: Agent modifies new requirement but acceptance looks at old; Runner tests commit A but Preview shows commit B; page says 'tests passed' but original command/output missing; human confirms a temporary page but release picks a different artifact. Therefore, the delivery platform's core storage is not a conclusion but an evidence mapping:
Delivery Evidence
├── requirement_revision
├── repository_identity
├── base_commit
├── candidate_commit
├── assignment_key
├── runner_identity
├── test_results
├── artifact_digest
├── evidence_signature
└── human_acceptanceThis mapping turns 'done' from a linguistic judgment into a verifiable relationship.
3. Why Requirement Confirmation and Task Dispatch Must Be in Same Reliable Process
An underestimated problem: requirement becomes 'confirmed' in DB, but the message to schedule Worker fails. If business transaction and message send are unrelated, you get: DB commit success + message send failure = page shows confirmed but no executor ever appears. ForgeX uses a transactional Outbox: confirming the requirement and registering the pending dispatch event happen in the same DB transaction; a scheduler later reliably delivers the event. Even if intermediate nodes briefly fail, the platform sees unsent events and continues processing, instead of letting tasks silently disappear. Such mechanisms don't make demo videos flashier, but determine whether the system works under real networks, process restarts, and concurrency.
4. Why Worker Needs Lease, Heartbeat, and Fencing Token
Worker picking up a task doesn't mean it owns execution rights forever. Machines disconnect, processes hang, tasks may be reassigned. The worst case: old Worker recovers and submits results while new Worker already completed the same task. ForgeX represents each execution right as a time-limited lease with an incrementing fencing token per assignment:
Old Worker: token = 17
Task times out and is reassigned
New Worker: token = 18
Old Worker recovers and submits → rejected due to expired tokenThe rejection targets the stale execution right, not the machine. Worker completion reports must carry task identity, project identity, repo identity, requirement version, base commit, candidate commit, branch, and assignment token; the platform verifies each field to prevent an execution result being incorrectly attached to another requirement.
5. Isolated Workspace Solves Delivery Contamination, Not Convenience
When a platform concurrently handles multiple requirements, all Agents cannot share a single writable directory. Git's official worktree mechanism allows a single repo to mount multiple working trees and check out different branches simultaneously. ForgeX leverages this isolation so each task runs in an independent workspace and branch, avoiding cross-contamination and preventing pollution of the developer's main workspace. After isolation, Worker's output should be a candidate commit — not a pile of 'files that look modified in the current directory'. The candidate commit establishes a critical invariant:
Subsequent Runner, Preview, and human acceptance must all point to the same unambiguous Git object.
6. Worker Completion ≠ Runner Pass
The phrase Coding Agents most love to say: 'Implementation complete, tests passed.' But the implementer's self-report is only a completion claim — it may contain useful info but cannot automatically become independent evidence. ForgeX separates Worker and Runner:
Worker changes code.
Runner receives only the specified repo and candidate commit.
Runner runs verification in an independent environment.
Verification results and artifact digests are written into evidence.
Platform verifies evidence signature and bindings before allowing state to advance.
In the end-to-end verification, Runner also generated a content-addressed Preview artifact: identified by SHA-256 digest of its content — any content change alters the digest. Evidence is signed with an independent key. The signature's purpose is not to prove 'code has no defects' but to prove: this evidence was produced by a trusted Runner for the specified object, and content was not silently swapped in transit or storage. This aligns with SLSA's core provenance idea: an artifact should not just have a download URL but be traceable to when, where, and how it was produced.
7. Preview Isn't an Ad-Hoc URL
Many teams treat Preview as 'spin up a machine and run the frontend'. But if the preview URL drifts over time or different users see different builds, it cannot serve acceptance. A delivery-grade Preview must: bind to the candidate commit, bind to the artifact digest, open in a restricted environment, not inherit unnecessary production permissions, and have the human acceptance action write back to the same delivery record. The end-to-end verification concluded with the Product role opening Preview, performing a real interaction, and confirming acceptance. Only then does the requirement enter 'completed' — which is still ForgeX delivery flow completion, not production release. This boundary must be preserved.
Chapter 3: From Verified Candidate to Real Production — There Must Be a Gate
1. 'Releasable' Is a Set of Checkable Conditions, Not a Status Word
Delivery stages can be split more precisely:
Worker completed : specified executor produced candidate commit; cannot prove independent verification passed.
Runner passed : specified commit passed agreed verification; cannot prove product experience meets expectations.
Human accepted : specified Preview accepted by business role; cannot prove production release authorized.
Release approved : change window, risk, owner confirmed; cannot prove production running successfully.
Deployed : specified artifact entered target environment; cannot prove business health, callbacks, data normal.
Production verified : post-deployment checks and observation done; cannot prove future issues won't arise.
A truly 'releasable' candidate must have: confirmed unambiguous requirement version; unique repo, branch, candidate commit; independent Runner verification on same commit; Preview artifact bound to commit with digest; explicit product/business acceptance record; reproducible build process and artifact provenance; clear change scope, release steps, rollback path; target environment, change window, release owner re-confirmed. The first five can be organized by ForgeX control plane; the last three typically integrate with existing CI/CD, artifact repository, change approval, release system, and ops.
2. Formal Release Should Continue Binding Same Object
The safest release publishes the already verified and accepted object — not 're-pull latest code and rebuild'. The ideal chain continues: Requirement Revision → Candidate Commit → Verified Artifact Digest → Release Approval → Deployment Record → Production Verification. Every step must use stable identities, not 'latest', 'current', or 'that package just now'. Immutable artifacts, digest verification, and provenance prevent drift in the last mile. Release also needs at least four safety mechanisms:
Explicit human gate : production actions must be authorized by someone with environment permissions and business responsibility. Low-risk reversible actions can be gradually automated; high-risk irreversible actions need stronger approval.
Recoverable changes : before release you must know how to roll back code, config, and data; 'we'll see if it fails' is not a rollback plan.
Progressive rollout : if you can canary, don't instantly replace all traffic. Observe health checks, key APIs, error rates, queues, callbacks, business metrics, then gradually expand.
Production evidence feedback : deployment success is not the endpoint. Production verification results should return to the same delivery record, closing the loop across requirement, commit, artifact, deployment, and runtime state.
3. What This Case Actually Proves
Combining the two evidence layers yields an honest conclusion table:
ForgeX can inventory and import source repos on a standard server: Yes (Bracelet real integration record).
ForgeX can identify current branch and untracked directories: Yes (Bracelet real workspace snapshot).
Project init checks Knowledge, Skill, MCP: Yes ( standard-delivery@1 init result).
Enters action_required and stops when conditions incomplete: Yes (Bracelet real state transition).
This integration did not directly modify Bracelet target source: Yes (integration scope and scene verification).
Worker actually modified Bracelet and produced candidate commit: No (never entered execution phase).
Runner verified Bracelet candidate commit: No (no candidate commit).
Bracelet Preview completed human acceptance: No (no Preview produced).
Full delivery control chain can go from requirement confirmation to human acceptance: Yes (real PostgreSQL end-to-end verification with simulated requirement).
Bracelet released to production: No (never concluded).
This table may not read like a smooth success story, but it gives engineering leads the real information: what's done, what's verified, what's blank. ForgeX itself once completed a release verification with DB migration, immutable release directory, service checks, and automated tests — a historical snapshot of platform release capability, not current real-time state, nor a substitute for Bracelet's never-happened release record. If we cannot distinguish 'real project evidence', 'automated verification evidence', and 'historical platform release evidence', we have no right to demand Agent accountability for delivery conclusions.
4. What Enterprises Should Optimize Isn't Agent Run Count
With multiple Agents, teams easily chase vanity metrics: daily Agent runs, lines of code generated, minutes to first commit, tasks without human intervention. These measure efficiency, not delivery quality. Better long-term metrics:
Project first-initialization pass rate. action_required mean time to close.
Requirement confirmation-to-dispatch success rate.
Expired lease and stale fencing token rejection count.
Candidate commit–Runner evidence full binding rate.
Preview–candidate artifact digest consistency rate.
Human acceptance rejection root causes.
Post-release rollback rate and production verification completion rate.
The former encourages Agents to do more; the latter encourages the system to do correctly, stop correctly, prove correctly.
5. What Is the Layer Above Agents That ForgeX Builds
Back to the series' starting point: why not just use Codex as a coding tool? Because coding is one execution action in the delivery chain. Above it, a company needs four coordinating planes:
Context Plane
Let Agent know real project, org, business context
MCP Permission & Audit Plane
Decide what capabilities this task may call, leave audit trail
Evidence Plane
Bind requirement, commit, test, artifact, signature, acceptance into verifiable evidence
Delivery Control Plane
Drive state transitions, schedule Worker and Runner, reliably stop when conditions unmetForgeX builds this delivery system above Agents. It doesn't replace Codex or compete with Coding Agents; it places Codex in a position with context, boundaries, evidence, independent verification, and human decision, letting model capability enter the company's real R&D process.
Conclusion
The most important result of this full-chain retrospective is not proving AI can auto-run from requirement to production. It confirms three fundamentals:
Real project integration: platform must respect the scene first.
Execution results must pass independent verification and bind to requirement version, code commit, Preview.
Between 'verified' and 'allowed to release', a clear human release gate must remain.
An enterprise AI R&D delivery platform shouldn't pursue having every requirement written by AI. It should guarantee: every requirement proceeds only when conditions met; stops reliably when not; every advance leaves evidence verifiable by the next responsible person. When that holds, Agent becomes a qualified executor in the company's delivery system, not just a faster code generator.
Fact Statement: Bracelet sections from real project integration record; full delivery chain from real PostgreSQL end-to-end verification with simulated business requirement; platform release data from historical verification snapshot. Three evidence types independent, none represent Bracelet auto-released to production.
Public References and Further Reading
Codex can be controlled programmatically as local Agent and integrated into internal tools and workflows.
Introduces workspace writes, network access, approval policies, sandbox boundaries.
Same repo can mount multiple worktrees and check out different branches simultaneously.
Explains how to trace artifact provenance with verifiable information: when, where, how produced.
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
