Why AI's 'Done' Isn't Delivery: The Evidence Plane Architecture
This article introduces the Evidence Plane architecture, arguing that AI agents' completion claims must be backed by independently verifiable evidence—including code diffs, test reports, security scans, approvals, and production metrics—organized through Evidence Contracts, automated verification, and traceable evidence bundles to enable trustworthy AI-driven software delivery.
Chapter 1: "Done" Is a Claim, Not a System State
1. Distinguishing Four Easily Confused Concepts
After an Agent task ends, the platform receives at least four distinct types of information:
Claim – the Agent's completion declaration (e.g., "Fixed login bug," "Added three tests," "All tests passed").
Transcript – the Agent's execution trace: analysis, tool calls, commands, file changes, errors, retries, approval requests. It shows what the Agent did, but does not prove the goal was achieved.
Outcome – the final state of the environment: whether the target change exists in the repo, the expected record is in the database, the page supports real business operations, the production service stays healthy.
Evidence – verifiable material supporting a conclusion, such as:
Acceptance criteria bound to the requirement version
Commit, diff, and build artifact digests
Reproducible test commands and structured reports
Playwright Trace, screenshots, network logs
Security scans, SBOM, artifact provenance
Approver identity, scope, and timestamp
Post-release logs, metrics, traces, and business results
The relationship can be summarized as:
Claim Agent claims what it completed
Transcript Agent actually executed what
Outcome System finally became what
Evidence Why we believe this resultOnly by separating these four can a company avoid mistaking an Agent's self-summary for system acceptance.
Figure 1: Left side shows Agent's 'done' claim; right side shows independent checks forming an evidence package. The diagram illustrates verification mechanisms, not that a specific model necessarily fabricates reports.
2. Can Agent-Written Tests Prove Correctness?
They can be part of the evidence but are not automatically sufficient. The Agent knows best what it changed, so having it add unit tests is reasonable. However, when developer and verifier share the same blind spots, tests may only cover scenarios the implementer already considered.
If a single Agent handles:
Understanding requirements
Writing implementation
Designing tests
Judging test coverage adequacy
Declaring readiness for release
the entire chain becomes self-proving.
This does not mean Agent-generated tests lack value; rather, we must distinguish:
Agent-generated tests → implementation evidence
Independent reviewer checks → review evidence
CI results in controlled environments → execution evidence
Product/business acceptance → goal evidence
Production metrics → real-result evidence
They prove different things and cannot substitute for each other.
3. "Tests Passed" Hides Too Much Information
A single "tests passed" statement may conceal:
Whether execution succeeded or only test collection completed
Whether one test file or the full regression suite ran
How many cases were skipped
Whether retries or flaky tests occurred
Whether the code version matches the final commit
Whether the runtime environment matches the target
Whether test data is realistic and valid
Whether tests cover the actual acceptance criteria
Therefore, the platform should store structured results, not just a natural-language summary:
command
exit_code
started_at
finished_at
source_revision
environment_digest
passed
failed
skipped
flaky
report_uri
artifact_digestFor browser tests, a final screenshot is often insufficient. Playwright Trace preserves action steps, DOM snapshots, console output, network requests, and errors, enabling re-examination of the failure process. Its value is not merely proving the page loaded, but letting verifiers step back through each action to observe what happened.
4. Why More Agents Make Evidence Problems More Acute
With one person using one Agent, context can be remembered and results manually checked. When a company runs dozens or hundreds of Agents in parallel:
Multiple Agents modify different modules concurrently
One Agent's output becomes another's input
The same task undergoes multiple failures and re-runs
Different environments use different configurations and permissions
Humans cannot read every execution trace line by line
The scarce resource shifts from coding time to verification capacity and human attention. OpenAI's Agent-first engineering emphasizes mechanically enforced architectural constraints, custom linters, and structural tests to maintain consistency because human review cannot keep pace with accelerated code generation. The Evidence Plane aims to organize scattered evidence from tasks, Agents, CI, tests, approvals, releases, and monitoring into a quality infrastructure that supports automated judgment, human spot-checks, and post-hoc traceability.
Chapter 2: Evidence Plane Is Not a Test Platform, But a Control Plane for Delivery Facts
1. Define the Evidence Contract First
Many teams ask "How should we verify?" after the Agent finishes. A safer approach is to define an Evidence Contract before execution. It answers: what evidence must appear for the task to be deemed complete, who produces it, what verifier is used, and what thresholds must be met.
Example Evidence Contract for a typical frontend requirement:
required:
- requirement_baseline
- acceptance_criteria
- source_diff
- lint_report
- typecheck_report
- unit_test_report
- e2e_test_report
- browser_trace_on_failure
- preview_url
- product_approval
gates:
lint: pass
typecheck: pass
unit_test_failed: 0
e2e_test_failed: 0
product_approval: requiredFor tasks involving database migrations, payments, permissions, or production config, the contract must also include:
Data backup and readability checks
Pre/post migration schema
Rollback scripts and trigger conditions
Security approvals
Canary results
Production business metrics
With completion criteria defined upfront, the Agent knows the target, and the verification system does not need to improvise based on the Agent's final claim.
2. Six Categories of Evidence the Evidence Plane Should Ingest
Intent Evidence – proves "why" and "what done means": requirement baseline, acceptance criteria, requirement changes, owner confirmation.
Change Evidence – proves "what actually changed": file diffs, commits, DB migrations, config changes, build inputs.
Verification Evidence – proves "what checks were run": lint, type checking, unit/integration/contract/E2E tests, screenshots, traces, manual exploration records.
Security & Supply Chain Evidence – secret scanning, SAST, SCA, container scanning, SBOM, artifact signing, provenance. SLSA defines provenance as verifiable information about when, where, and how a software artifact was produced, enabling traceability to source and build process.
Decision Evidence – records who approved/rejected, in what role, scope, and risk conditions. Approval must store the exact version of the approved object, not just "agreed".
Production Evidence – proves "what happened after release": health checks, error rates, latency, resource metrics, business success rates, alerts, rollbacks, incident records.
3. A Practical Six-Layer Evidence Plane Architecture
Evidence Producers – requirement systems, Git, Agent Runtime, CI, test platforms, security tools, approval systems, release platforms, observability systems.
Evidence Collectors – receive structured events, test reports, file artifacts, external references; enrich with task, project, commit, Agent run, and environment identity.
Evidence Registry – stores evidence metadata, summaries, provenance, scope, and relationships. Large traces, videos, build artifacts live in object storage but must have stable URIs and content digests in the Registry.
Verifiers – check evidence authenticity, completeness, correct versioning; execute deterministic rules or controlled scoring.
Quality Gates – aggregate judgments per task risk and Evidence Contract; output PASS, FAIL, ACTION_REQUIRED, or APPROVAL_REQUIRED.
Audit & Feedback – feed missing evidence, human rejections, production defects, and incidents back into requirements, Skills, test suites, and quality rules.
Figure 2: Evidence Plane conceptual architecture. Requirements, code, build, test, security, approval, and production telemetry produce evidence; registry stores provenance and links; independent verifiers judge per policy. The diagram does not prescribe specific products or tech stacks.
4. Core Data Model: What to Persist
A base evidence model includes:
evidence
├── id
├── tenant_id
├── project_id
├── requirement_id
├── task_id
├── stage_id
├── agent_run_id
├── evidence_type
├── producer_type
├── producer_identity
├── source_revision
├── environment
├── uri
├── sha256
├── status
├── summary
├── created_at
└── expires_atBut a single evidence item rarely decides delivery. Two additional objects are needed:
evidence_bundle – groups all evidence for a candidate version: requirements, changes, tests, security, approvals, production evidence.
gate_decision – records:
Evidence Contract version used
Which evidence was checked
Which evidence was missing
Each verifier's result
Final decision
Decision timestamp
Human exceptions and reasons
This enables the platform to answer:
Why was this version allowed to proceed to the next stage?
5. Evidence Must Bind to Precise Objects, Not Just Task Names
Suppose a task goes through three Agent runs: first test fails, second passes after a fix, third adjusts another file. If the platform only shows "task tests passed", it might incorrectly apply the second run's test results to the third code version. Therefore, evidence must bind at least to:
Requirement Version
Task Attempt
Agent Run
Source Revision
Build Artifact Digest
Target EnvironmentAny change to a key object should trigger gate re-evaluation. Old results are not evidence for new versions.
6. Evidence Must Have Provenance and Strength
Evidence trustworthiness varies:
Agent natural-language summary
↓
Unstructured command output
↓
Controlled runner generated reports
↓
CI platform commit-bound artifacts
↓
Immutable artifacts with digest, signature, provenance
↓
Independent environment result verification & production observationThis is not a strict replacement hierarchy; rather, evidence closer to the real environment, more independent from the implementer, and harder to tamper with post-hoc is better suited for high-risk decisions. Evidence must also record its scope of applicability: a unit test suite proves certain function behaviors but not the full business flow; a screenshot proves a screen appeared once but not that buttons, APIs, and permissions work. A rigorous Evidence Plane does not promise "absolute safety"; it tells decision-makers exactly which evidence supports the current conclusion and which risks remain uncovered.
Chapter 3: Drive Agent Delivery with Evidence, Not Trust
1. Verification Must Match Risk
Not all tasks need the same weight of evidence. A copy change and a production database migration should not pass the same gate. Three risk-based baselines:
Low-risk changes : Diff + linting + basic tests + human spot-check
Medium-risk changes : Requirement baseline + independent review + full CI + integration/E2E + security scans + preview acceptance
High-risk changes : All medium-risk evidence + backup/rollback verification + multi-role approval + immutable artifacts with provenance + canary + production metric comparison + automated or manual rollback decision
Real efficiency is not removing verification, but avoiding forcing every task through the highest verification cost while ensuring high-risk tasks cannot pass on a mere "looks fine".
2. Separate Implementer, Verifier, and Approver
The Agent era still requires separation of duties, though participants may be humans and different Agents:
Developer Agent responsible for implementation
Reviewer Agent responsible for independent review
Runner responsible for deterministic execution
Test Agent responsible for designing tests from acceptance criteria
Security Tools responsible for security & supply chain checks
Human Approver responsible for risk, exceptions, accountabilityThe key is that the verification path must not rely entirely on the same implementation context. Independent verifiers see requirements, diffs, and the Evidence Contract but need not inherit the developer Agent's full reasoning; context isolation reduces bias. Model-based Graders suit language quality, requirement conformance, and subjective criteria; code-based Graders suit tests, static analysis, state checks; humans calibrate standards, handle exceptions, and bear final responsibility. Anthropic's Agent Eval practice also stresses combining code-based, model-based, and human evaluation because a single evaluation layer cannot cover all failure modes.
3. The Evidence Chain Continues After Release
Deployment success only means the release action completed, not that business results are correct. Post-release, continuous collection is needed:
Service health
Error rate and latency anomalies
Core business success rate drops
New version triggered alerts
User paths uncovered by prior tests
OpenTelemetry provides traces, metrics, and logs as complementary signals. The Evidence Plane's value is linking these runtime signals back to Release, Commit, Requirement, and Agent Run. For example:
Business success rate drops
↓
Release 2.8.31
↓
Build Artifact sha256:...
↓
Commit abc123
↓
Agent Run #912
↓
Task #327
↓
Requirement #102Thus, a production anomaly becomes an automatic trace back through the full delivery chain.
Figure 3: Evidence-driven delivery loop. Implementation Agent and verification stages are separated; evidence accumulates along requirements, code, test, security, approval, artifact, release, and production observation, feeding back into the next improvement cycle.
4. ForgeX Should Show an Evidence Map, Not a "Success" Message
In ForgeX, a requirement should not end with just: Status: Completed A more valuable page is an Evidence Map :
REQ-102 User Login Experience Optimization
Requirement Baseline ✅ v3
Code Change ✅ Commit abc123
Agent Run ✅ Run #912
Independent Review ✅ 0 Blocker
Lint ✅
TypeCheck ✅
Unit Test ✅ 328 passed / 0 failed / 4 skipped
E2E ✅ 12 passed / Trace available
Security ✅ Critical 0 / High 0
Preview ✅ URL + screenshot
Product Acceptance ✅ Li Si / 2026-08-14
Release Artifact ✅ sha256:...
Canary ✅
Production ✅ Error Rate / P95 / Business KPI normalIf key evidence is missing, show ACTION_REQUIRED; if evidence fails, show FAIL; if auto-checks pass but human risk decision remains, show APPROVAL_REQUIRED. This is more meaningful than collapsing all states into "In Progress" and "Completed". ForgeX's actual delivery verification already follows this direction: requirements go through formal confirmation, Worker, Runner, preview, and deployment chains, reporting end-to-end evidence rather than bypassing the platform to modify target projects directly. Specific version numbers, test counts, and runtime environments are historical state; public documentation should re-collect from the current version.
5. What Management Should Really Watch
With an Evidence Plane, management need not read every Agent dialog. More valuable metrics include:
Evidence Completeness : Evidence Contract fulfillment rate, missing evidence count
First-Time Quality : First gate pass rate, first acceptance pass rate
Human Attention : Per-task manual review time, human intervention ratio
Verification Effectiveness : Production escape defects, false pass rate, false block rate
Delivery Efficiency : Cycle from requirement confirmation to evidence completeness
Traceability : Proportion of releases traceable to requirements and Agent runs
Production Quality : Canary failures, rollbacks, incidents, business metric anomalies
These metrics don't directly tell "which model is best", but answer a more important question:
Can we reliably turn Agent output into trustworthy delivery?
Conclusion: In the AI Era, the Scarcest Resource Is Not Generation Capability, But Trustworthy Completion
As the cost of generating code, docs, and designs keeps falling, organizations easily accumulate more and more "looks done" results. But software delivery has never needed more completion claims.
Companies need to know:
Whether requirements are truly satisfied
Whether code comes from the correct version
Whether checks ran in the correct environment
What tests covered and what they missed
Who approved which risks
Whether users got expected results after release
Whether issues can be traced and rolled back
That is why the Evidence Plane becomes the core infrastructure of company-level AI R&D delivery platforms. It is not to prove AI untrustworthy, nor to add more approvals to every step. Its true goal is to shift trust from "believing an Agent's statement" to "believing an evidence chain that has provenance, scope, versioning, independent verification, and accountability".
In the future, Agents will take on more execution work. Human roles will shift toward defining success, designing verification, approving exceptions, and bearing responsibility. When a company can clearly distinguish Claim, Transcript, Outcome, and Evidence; define Evidence Contracts before tasks; and continue verifying real results after release, it possesses the foundation for large-scale AI Agent adoption.
So the next time AI tells you "done", the most important question is not:
Are you sure?
But:
Where is the evidence, and what exactly does it prove?
References & Scope Notes
Anthropic – referenced for distinctions among Task, Trial, Grader, Transcript, Outcome, Evaluation Harness, and the idea of combining code-based, model-based, and human evaluation.
OpenAI – referenced for Agent-first engineering practices using structural tests, custom linters, architectural constraints, and feedback loops; the case does not imply all organizations achieve the same results.
OpenAI – referenced for the principle of "first clarify the conclusion to be supported, then describe the evaluation apparatus and valid evidence".
SLSA – referenced for artifact provenance, verifiable build information, and attestation concepts; the article does not claim its conceptual architecture automatically satisfies any SLSA level.
Playwright – referenced for Trace capabilities (actions, DOM snapshots, logs, network, errors); whether to mandate it as required evidence depends on task risk and cost.
OpenTelemetry – referenced for traces, metrics, logs as post-release evidence sources; telemetry existence does not automatically prove business correctness, still requires user-outcome-oriented metrics.
Evidence Contract, Evidence Registry, Evidence Bundle, Gate Decision, Evidence Map – concepts proposed in this article for company-level AI delivery models. Real systems must adapt to industry regulation, data sensitivity, existing CI/CD, approval processes, and production risks.
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
