R&D Management 31 min read

Why AI's 'Done' Isn't Delivery: The Evidence Plane Architecture

This article introduces the Evidence Plane architecture, arguing that AI agents' completion claims must be backed by independently verifiable evidence—including code diffs, test reports, security scans, approvals, and production metrics—organized through Evidence Contracts, automated verification, and traceable evidence bundles to enable trustworthy AI-driven software delivery.

Chengwu Tech Stack
Chengwu Tech Stack
Chengwu Tech Stack
Why AI's 'Done' Isn't Delivery: The Evidence Plane Architecture

Chapter 1: "Done" Is a Claim, Not a System State

1. Distinguishing Four Easily Confused Concepts

After an Agent task ends, the platform receives at least four distinct types of information:

Claim – the Agent's completion declaration (e.g., "Fixed login bug," "Added three tests," "All tests passed").

Transcript – the Agent's execution trace: analysis, tool calls, commands, file changes, errors, retries, approval requests. It shows what the Agent did, but does not prove the goal was achieved.

Outcome – the final state of the environment: whether the target change exists in the repo, the expected record is in the database, the page supports real business operations, the production service stays healthy.

Evidence – verifiable material supporting a conclusion, such as:

Acceptance criteria bound to the requirement version

Commit, diff, and build artifact digests

Reproducible test commands and structured reports

Playwright Trace, screenshots, network logs

Security scans, SBOM, artifact provenance

Approver identity, scope, and timestamp

Post-release logs, metrics, traces, and business results

The relationship can be summarized as:

Claim        Agent claims what it completed
Transcript   Agent actually executed what
Outcome      System finally became what
Evidence     Why we believe this result

Only by separating these four can a company avoid mistaking an Agent's self-summary for system acceptance.

Figure 1: Left side shows Agent's 'done' claim; right side shows independent checks forming an evidence package. The diagram illustrates verification mechanisms, not that a specific model necessarily fabricates reports.
Figure 1: Left side shows Agent's 'done' claim; right side shows independent checks forming an evidence package. The diagram illustrates verification mechanisms, not that a specific model necessarily fabricates reports.

Figure 1: Left side shows Agent's 'done' claim; right side shows independent checks forming an evidence package. The diagram illustrates verification mechanisms, not that a specific model necessarily fabricates reports.

2. Can Agent-Written Tests Prove Correctness?

They can be part of the evidence but are not automatically sufficient. The Agent knows best what it changed, so having it add unit tests is reasonable. However, when developer and verifier share the same blind spots, tests may only cover scenarios the implementer already considered.

If a single Agent handles:

Understanding requirements

Writing implementation

Designing tests

Judging test coverage adequacy

Declaring readiness for release

the entire chain becomes self-proving.

This does not mean Agent-generated tests lack value; rather, we must distinguish:

Agent-generated tests → implementation evidence

Independent reviewer checks → review evidence

CI results in controlled environments → execution evidence

Product/business acceptance → goal evidence

Production metrics → real-result evidence

They prove different things and cannot substitute for each other.

3. "Tests Passed" Hides Too Much Information

A single "tests passed" statement may conceal:

Whether execution succeeded or only test collection completed

Whether one test file or the full regression suite ran

How many cases were skipped

Whether retries or flaky tests occurred

Whether the code version matches the final commit

Whether the runtime environment matches the target

Whether test data is realistic and valid

Whether tests cover the actual acceptance criteria

Therefore, the platform should store structured results, not just a natural-language summary:

command
exit_code
started_at
finished_at
source_revision
environment_digest
passed
failed
skipped
flaky
report_uri
artifact_digest

For browser tests, a final screenshot is often insufficient. Playwright Trace preserves action steps, DOM snapshots, console output, network requests, and errors, enabling re-examination of the failure process. Its value is not merely proving the page loaded, but letting verifiers step back through each action to observe what happened.

4. Why More Agents Make Evidence Problems More Acute

With one person using one Agent, context can be remembered and results manually checked. When a company runs dozens or hundreds of Agents in parallel:

Multiple Agents modify different modules concurrently

One Agent's output becomes another's input

The same task undergoes multiple failures and re-runs

Different environments use different configurations and permissions

Humans cannot read every execution trace line by line

The scarce resource shifts from coding time to verification capacity and human attention. OpenAI's Agent-first engineering emphasizes mechanically enforced architectural constraints, custom linters, and structural tests to maintain consistency because human review cannot keep pace with accelerated code generation. The Evidence Plane aims to organize scattered evidence from tasks, Agents, CI, tests, approvals, releases, and monitoring into a quality infrastructure that supports automated judgment, human spot-checks, and post-hoc traceability.

Chapter 2: Evidence Plane Is Not a Test Platform, But a Control Plane for Delivery Facts

1. Define the Evidence Contract First

Many teams ask "How should we verify?" after the Agent finishes. A safer approach is to define an Evidence Contract before execution. It answers: what evidence must appear for the task to be deemed complete, who produces it, what verifier is used, and what thresholds must be met.

Example Evidence Contract for a typical frontend requirement:

required:
  - requirement_baseline
  - acceptance_criteria
  - source_diff
  - lint_report
  - typecheck_report
  - unit_test_report
  - e2e_test_report
  - browser_trace_on_failure
  - preview_url
  - product_approval

gates:
  lint: pass
  typecheck: pass
  unit_test_failed: 0
  e2e_test_failed: 0
  product_approval: required

For tasks involving database migrations, payments, permissions, or production config, the contract must also include:

Data backup and readability checks

Pre/post migration schema

Rollback scripts and trigger conditions

Security approvals

Canary results

Production business metrics

With completion criteria defined upfront, the Agent knows the target, and the verification system does not need to improvise based on the Agent's final claim.

2. Six Categories of Evidence the Evidence Plane Should Ingest

Intent Evidence – proves "why" and "what done means": requirement baseline, acceptance criteria, requirement changes, owner confirmation.

Change Evidence – proves "what actually changed": file diffs, commits, DB migrations, config changes, build inputs.

Verification Evidence – proves "what checks were run": lint, type checking, unit/integration/contract/E2E tests, screenshots, traces, manual exploration records.

Security & Supply Chain Evidence – secret scanning, SAST, SCA, container scanning, SBOM, artifact signing, provenance. SLSA defines provenance as verifiable information about when, where, and how a software artifact was produced, enabling traceability to source and build process.

Decision Evidence – records who approved/rejected, in what role, scope, and risk conditions. Approval must store the exact version of the approved object, not just "agreed".

Production Evidence – proves "what happened after release": health checks, error rates, latency, resource metrics, business success rates, alerts, rollbacks, incident records.

3. A Practical Six-Layer Evidence Plane Architecture

Evidence Producers – requirement systems, Git, Agent Runtime, CI, test platforms, security tools, approval systems, release platforms, observability systems.

Evidence Collectors – receive structured events, test reports, file artifacts, external references; enrich with task, project, commit, Agent run, and environment identity.

Evidence Registry – stores evidence metadata, summaries, provenance, scope, and relationships. Large traces, videos, build artifacts live in object storage but must have stable URIs and content digests in the Registry.

Verifiers – check evidence authenticity, completeness, correct versioning; execute deterministic rules or controlled scoring.

Quality Gates – aggregate judgments per task risk and Evidence Contract; output PASS, FAIL, ACTION_REQUIRED, or APPROVAL_REQUIRED.

Audit & Feedback – feed missing evidence, human rejections, production defects, and incidents back into requirements, Skills, test suites, and quality rules.

Figure 2: Evidence Plane conceptual architecture. Requirements, code, build, test, security, approval, and production telemetry produce evidence; registry stores provenance and links; independent verifiers judge per policy. The diagram does not prescribe specific products or tech stacks.
Figure 2: Evidence Plane conceptual architecture. Requirements, code, build, test, security, approval, and production telemetry produce evidence; registry stores provenance and links; independent verifiers judge per policy. The diagram does not prescribe specific products or tech stacks.

Figure 2: Evidence Plane conceptual architecture. Requirements, code, build, test, security, approval, and production telemetry produce evidence; registry stores provenance and links; independent verifiers judge per policy. The diagram does not prescribe specific products or tech stacks.

4. Core Data Model: What to Persist

A base evidence model includes:

evidence
├── id
├── tenant_id
├── project_id
├── requirement_id
├── task_id
├── stage_id
├── agent_run_id
├── evidence_type
├── producer_type
├── producer_identity
├── source_revision
├── environment
├── uri
├── sha256
├── status
├── summary
├── created_at
└── expires_at

But a single evidence item rarely decides delivery. Two additional objects are needed:

evidence_bundle – groups all evidence for a candidate version: requirements, changes, tests, security, approvals, production evidence.

gate_decision – records:

Evidence Contract version used

Which evidence was checked

Which evidence was missing

Each verifier's result

Final decision

Decision timestamp

Human exceptions and reasons

This enables the platform to answer:

Why was this version allowed to proceed to the next stage?

5. Evidence Must Bind to Precise Objects, Not Just Task Names

Suppose a task goes through three Agent runs: first test fails, second passes after a fix, third adjusts another file. If the platform only shows "task tests passed", it might incorrectly apply the second run's test results to the third code version. Therefore, evidence must bind at least to:

Requirement Version
Task Attempt
Agent Run
Source Revision
Build Artifact Digest
Target Environment

Any change to a key object should trigger gate re-evaluation. Old results are not evidence for new versions.

6. Evidence Must Have Provenance and Strength

Evidence trustworthiness varies:

Agent natural-language summary
      ↓
Unstructured command output
      ↓
Controlled runner generated reports
      ↓
CI platform commit-bound artifacts
      ↓
Immutable artifacts with digest, signature, provenance
      ↓
Independent environment result verification & production observation

This is not a strict replacement hierarchy; rather, evidence closer to the real environment, more independent from the implementer, and harder to tamper with post-hoc is better suited for high-risk decisions. Evidence must also record its scope of applicability: a unit test suite proves certain function behaviors but not the full business flow; a screenshot proves a screen appeared once but not that buttons, APIs, and permissions work. A rigorous Evidence Plane does not promise "absolute safety"; it tells decision-makers exactly which evidence supports the current conclusion and which risks remain uncovered.

Chapter 3: Drive Agent Delivery with Evidence, Not Trust

1. Verification Must Match Risk

Not all tasks need the same weight of evidence. A copy change and a production database migration should not pass the same gate. Three risk-based baselines:

Low-risk changes : Diff + linting + basic tests + human spot-check

Medium-risk changes : Requirement baseline + independent review + full CI + integration/E2E + security scans + preview acceptance

High-risk changes : All medium-risk evidence + backup/rollback verification + multi-role approval + immutable artifacts with provenance + canary + production metric comparison + automated or manual rollback decision

Real efficiency is not removing verification, but avoiding forcing every task through the highest verification cost while ensuring high-risk tasks cannot pass on a mere "looks fine".

2. Separate Implementer, Verifier, and Approver

The Agent era still requires separation of duties, though participants may be humans and different Agents:

Developer Agent      responsible for implementation
Reviewer Agent       responsible for independent review
Runner               responsible for deterministic execution
Test Agent           responsible for designing tests from acceptance criteria
Security Tools       responsible for security & supply chain checks
Human Approver       responsible for risk, exceptions, accountability

The key is that the verification path must not rely entirely on the same implementation context. Independent verifiers see requirements, diffs, and the Evidence Contract but need not inherit the developer Agent's full reasoning; context isolation reduces bias. Model-based Graders suit language quality, requirement conformance, and subjective criteria; code-based Graders suit tests, static analysis, state checks; humans calibrate standards, handle exceptions, and bear final responsibility. Anthropic's Agent Eval practice also stresses combining code-based, model-based, and human evaluation because a single evaluation layer cannot cover all failure modes.

3. The Evidence Chain Continues After Release

Deployment success only means the release action completed, not that business results are correct. Post-release, continuous collection is needed:

Service health

Error rate and latency anomalies

Core business success rate drops

New version triggered alerts

User paths uncovered by prior tests

OpenTelemetry provides traces, metrics, and logs as complementary signals. The Evidence Plane's value is linking these runtime signals back to Release, Commit, Requirement, and Agent Run. For example:

Business success rate drops
  ↓
Release 2.8.31
  ↓
Build Artifact sha256:...
  ↓
Commit abc123
  ↓
Agent Run #912
  ↓
Task #327
  ↓
Requirement #102

Thus, a production anomaly becomes an automatic trace back through the full delivery chain.

Figure 3: Evidence-driven delivery loop. Implementation Agent and verification stages are separated; evidence accumulates along requirements, code, test, security, approval, artifact, release, and production observation, feeding back into the next improvement cycle.
Figure 3: Evidence-driven delivery loop. Implementation Agent and verification stages are separated; evidence accumulates along requirements, code, test, security, approval, artifact, release, and production observation, feeding back into the next improvement cycle.

Figure 3: Evidence-driven delivery loop. Implementation Agent and verification stages are separated; evidence accumulates along requirements, code, test, security, approval, artifact, release, and production observation, feeding back into the next improvement cycle.

4. ForgeX Should Show an Evidence Map, Not a "Success" Message

In ForgeX, a requirement should not end with just: Status: Completed A more valuable page is an Evidence Map :

REQ-102  User Login Experience Optimization

Requirement Baseline     ✅ v3
Code Change              ✅ Commit abc123
Agent Run                ✅ Run #912
Independent Review       ✅ 0 Blocker
Lint                     ✅
TypeCheck                ✅
Unit Test                ✅ 328 passed / 0 failed / 4 skipped
E2E                      ✅ 12 passed / Trace available
Security                 ✅ Critical 0 / High 0
Preview                  ✅ URL + screenshot
Product Acceptance       ✅ Li Si / 2026-08-14
Release Artifact         ✅ sha256:...
Canary                   ✅
Production               ✅ Error Rate / P95 / Business KPI normal

If key evidence is missing, show ACTION_REQUIRED; if evidence fails, show FAIL; if auto-checks pass but human risk decision remains, show APPROVAL_REQUIRED. This is more meaningful than collapsing all states into "In Progress" and "Completed". ForgeX's actual delivery verification already follows this direction: requirements go through formal confirmation, Worker, Runner, preview, and deployment chains, reporting end-to-end evidence rather than bypassing the platform to modify target projects directly. Specific version numbers, test counts, and runtime environments are historical state; public documentation should re-collect from the current version.

5. What Management Should Really Watch

With an Evidence Plane, management need not read every Agent dialog. More valuable metrics include:

Evidence Completeness : Evidence Contract fulfillment rate, missing evidence count

First-Time Quality : First gate pass rate, first acceptance pass rate

Human Attention : Per-task manual review time, human intervention ratio

Verification Effectiveness : Production escape defects, false pass rate, false block rate

Delivery Efficiency : Cycle from requirement confirmation to evidence completeness

Traceability : Proportion of releases traceable to requirements and Agent runs

Production Quality : Canary failures, rollbacks, incidents, business metric anomalies

These metrics don't directly tell "which model is best", but answer a more important question:

Can we reliably turn Agent output into trustworthy delivery?

Conclusion: In the AI Era, the Scarcest Resource Is Not Generation Capability, But Trustworthy Completion

As the cost of generating code, docs, and designs keeps falling, organizations easily accumulate more and more "looks done" results. But software delivery has never needed more completion claims.

Companies need to know:

Whether requirements are truly satisfied

Whether code comes from the correct version

Whether checks ran in the correct environment

What tests covered and what they missed

Who approved which risks

Whether users got expected results after release

Whether issues can be traced and rolled back

That is why the Evidence Plane becomes the core infrastructure of company-level AI R&D delivery platforms. It is not to prove AI untrustworthy, nor to add more approvals to every step. Its true goal is to shift trust from "believing an Agent's statement" to "believing an evidence chain that has provenance, scope, versioning, independent verification, and accountability".

In the future, Agents will take on more execution work. Human roles will shift toward defining success, designing verification, approving exceptions, and bearing responsibility. When a company can clearly distinguish Claim, Transcript, Outcome, and Evidence; define Evidence Contracts before tasks; and continue verifying real results after release, it possesses the foundation for large-scale AI Agent adoption.

So the next time AI tells you "done", the most important question is not:

Are you sure?

But:

Where is the evidence, and what exactly does it prove?

References & Scope Notes

Anthropic – referenced for distinctions among Task, Trial, Grader, Transcript, Outcome, Evaluation Harness, and the idea of combining code-based, model-based, and human evaluation.

OpenAI – referenced for Agent-first engineering practices using structural tests, custom linters, architectural constraints, and feedback loops; the case does not imply all organizations achieve the same results.

OpenAI – referenced for the principle of "first clarify the conclusion to be supported, then describe the evaluation apparatus and valid evidence".

SLSA – referenced for artifact provenance, verifiable build information, and attestation concepts; the article does not claim its conceptual architecture automatically satisfies any SLSA level.

Playwright – referenced for Trace capabilities (actions, DOM snapshots, logs, network, errors); whether to mandate it as required evidence depends on task risk and cost.

OpenTelemetry – referenced for traces, metrics, logs as post-release evidence sources; telemetry existence does not automatically prove business correctness, still requires user-outcome-oriented metrics.

Evidence Contract, Evidence Registry, Evidence Bundle, Gate Decision, Evidence Map – concepts proposed in this article for company-level AI delivery models. Real systems must adapt to industry regulation, data sensitivity, existing CI/CD, approval processes, and production risks.

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

AI agentsquality gatesTraceabilitysoftware deliveryevidence-based verificationEvidence ContractEvidence Planerisk-based verification
Chengwu Tech Stack
Written by

Chengwu Tech Stack

A powerful mindset is a lifelong treasure!

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.