AI-Native R&D OS Architecture: Codex, Temporal & MCP for End-to-End Delivery
This article details the technical architecture of a company-level AI R&D delivery platform V1.0, positioning it as an AI-native operating system that orchestrates Codex agents via Temporal workflows, connects enterprise systems through an MCP gateway, enforces quality via deterministic CI gates, and ensures full traceability from requirements to production incidents.
Platform Positioning
The platform is not a Codex wrapper nor a traditional Jira/Zentao with an AI chat box. It is an AI-native R&D delivery platform centered on projects and requirements, using Temporal workflows for process control, Codex as the execution engine, MCP as the enterprise system connectivity layer, Skills as the company's R&D methodology, and human approvals plus deterministic CI for quality assurance.
All roles (product, development, test, ops) interact with a unified platform that sits above a stage workflow engine, an Agent Orchestrator, specialized agents (Product, Codex Developer, Test/Ops), a Skills Registry, and an MCP Gateway connecting to Plane, GitLab, CI, Database, ArgoCD, Grafana/OTel.
V1.0 Recommended Tech Stack
Company Web Portal : Custom React/Next.js
Project/Task Base : Plane Self-hosted (using Work Item API; old /issues/ API deprecated after 2026-03-31)
Workflow Engine : Temporal
AI Execution : Codex SDK + App Server
MCP Gateway : IBM ContextForge
Skills : Git + Codex Skills
Auth : Keycloak
Authorization : Platform RBAC + ABAC
Database : PostgreSQL
Cache : Redis
Object Storage : MinIO/S3
Code Platform : GitLab Self-Managed
CI : GitLab CI
E2E Testing : Playwright
CD : Argo CD
Security Management : DefectDojo + Dependency-Track
Observability : OpenTelemetry + Grafana
Agent Observability : Langfuse
Container Runtime : Kubernetes
Avoid Deep Forking of Plane
Do not fork Plane and modify tens of thousands of lines; upgrades become painful. Instead, build a Platform API with a Plane Adapter that uses Plane's REST API and Webhooks. Plane handles Project, Work Item, Module, Cycle, basic comments, Wiki/Docs, and ordinary task assignment. The platform owns Requirement Baseline, Stage Status, Agent, Skills, MCP, Approval, Quality Gate, Evidence, Release, Ops, and Audit.
Backend Architecture: One Control Plane + Three Worker Types
Avoid dozens of microservices. Structure:
ai-rd-platform/
├── platform-web
├── platform-api
│ ├── project
│ ├── requirement
│ ├── workflow
│ ├── approval
│ ├── agent
│ ├── evidence
│ ├── quality
│ ├── release
│ ├── incident
│ └── audit
├── workflow-worker
├── agent-orchestrator
└── codex-executorOnly Codex Executor × N , Test Worker × N , and Workflow Worker × N need horizontal scaling.
Core Business Data Models
Plane Stores
Workspace, Project, WorkItem, Module, Cycle, Comment, DocumentPlatform PostgreSQL Stores
1. requirement_baseline
id, project_id, plane_work_item_id, version, content_snapshot, acceptance_criteria, created_by, approved_by, approved_at, statusExample: REQ-102 V1 → Product Approval → Frozen. Development agents must read the frozen baseline, never the latest chat.
2. workflow_instance
id, project_id, requirement_id, temporal_workflow_id, current_stage, status, created_at, completed_at3. stage_instance
id, workflow_id, stage_type, status (INPUT_READY, AI_RUNNING, WAIT_REVIEW, REJECTED, APPROVED, COMPLETED)4. approval
id, stage_id, approval_type (PRODUCT_APPROVAL, TECH_APPROVAL, QA_APPROVAL, OPS_APPROVAL, RELEASE_APPROVAL, SECURITY_EXCEPTION), approver_user_id, decision, comment, created_at5. agent_task
id, workflow_id, stage_id, agent_type, requirement_version, repository, branch, skill_profile, mcp_profile, status, priority, created_by6. agent_run
run_id, agent_task_id, codex_thread_id, model, started_at, finished_at, input_context_hash, skill_versions, mcp_policy_version, git_before_sha, git_after_sha, status, token_input, token_output, costCritical for traceability: which agent, which requirement version, which skill version generated the code.
Evidence: Core of Black-Box Development Traceability
Product managers and leaders need not read code, but the platform must store machine execution evidence .
evidence
id, project_id, requirement_id, stage_id, agent_run_id, evidence_type, uri, sha256, producer, created_atEvidence types: PRD, ARCHITECTURE, ADR, SOURCE_DIFF, BUILD_LOG, UNIT_TEST_REPORT, INTEGRATION_TEST_REPORT, E2E_REPORT, E2E_SCREENSHOT, PLAYWRIGHT_TRACE, COVERAGE_REPORT, SAST_REPORT, SCA_REPORT, SBOM, RELEASE_MANIFEST, DEPLOYMENT_LOG, PRODUCTION_METRICS, INCIDENT_REPORT.
This creates a full chain: REQ-102 → PRD V3 → Architecture V2 → TASK-332 → Codex Run #873 → MR !932 → Pipeline #1827 → Unit Test → E2E Trace → Security Report → Release 2.3.1 → Production Verification.
Temporal as Company-Wide R&D State Machine
Replace manual status enums (if status == 17) with a Temporal workflow covering the entire lifecycle:
DeliveryWorkflow
├── RequirementAnalysis
├── ProductApproval
├── TechnicalDesign
├── TechnicalApproval
├── Development
├── CodeReview
├── QA
├── ProductAcceptance
├── ReleaseApproval
├── Deployment
├── ProductionVerification
└── OperationsTemporal's strength: workflows survive process crashes, server failures, and long waits (e.g., AI execution → wait 2 days for human approval → continue AI execution).
Complete Requirement Workflow States
NEW → REQUIREMENT_ANALYSIS → WAIT_PRODUCT_APPROVAL → REQUIREMENT_APPROVED → TECH_DESIGN → WAIT_TECH_APPROVAL → TECH_APPROVED → TASK_DECOMPOSITION → DEVELOPMENT → CODE_REVIEW → CI_GATE → QA_TESTING → WAIT_QA_APPROVAL → PRODUCT_ACCEPTANCE → WAIT_PRODUCT_ACCEPTANCE → RELEASE_PREPARATION → WAIT_RELEASE_APPROVAL → CANARY_DEPLOYMENT → PRODUCTION_VERIFICATION → RELEASED → OPERATIONSAny stage can: Approve, Reject, Return, Cancel, Retry, Escalate.
Parallel Codex Agents in Development Phase
Example: REQ-102 → Task Decomposition Agent → parallel Backend, Web, iOS, Android Agents each with independent container + Git worktree + agent identity + scoped MCP + scoped Skills → Integration Agent → CI. Forbid multiple agents sharing a work directory.
Codex Execution Service (Node.js/TypeScript)
Use official Codex SDK (TypeScript) for server-side thread start/continue/resume. Structure:
Temporal → Agent Orchestrator → Codex Execution Service
├── create workspace
├── git clone/fetch
├── git worktree
├── load AGENTS.md
├── mount Skills
├── create MCP credentials
├── start Codex thread
├── stream events
├── collect diff
├── run verification
├── create MR
└── destroy workspaceCodex App Server vs SDK
SDK : Unattended tasks, CI, auto-dev, auto-test, background agents.
App Server : Human watching agent work (real-time UI with plan, current file diff, command approval prompts).
Thus: Background auto-execution → Codex SDK; Foreground real-time Agent UI → Codex App Server.
Agent Orchestrator: The AI Layer Brain
Does not write code. Responsibilities:
Select Agent → Select Skill → Select MCP Permissions → Construct Context → Launch Codex → Receive Result → Judge Gate Entry → Return to TemporalAgent Profile example (backend-developer):
name: backend-developer
skills:
- requirement-read
- impact-analysis
- backend-development
- unit-test
mcp_profile:
- plane-read
- git-write
- ci-trigger
- db-test
environment:
production: deny
output:
- implementation
- unit_tests
- change_summary
- risk_summaryCompany Agent Catalog
At minimum: ProductRequirementAgent, ArchitectureAgent, TaskPlannerAgent, BackendDeveloperAgent, FrontendDeveloperAgent, MobileDeveloperAgent, IntegrationAgent, CodeReviewerAgent, TestDesignerAgent, TestExecutionAgent, SecurityReviewAgent, ReleaseAgent, OperationsAgent, IncidentAgent.
Critical isolation : Developer Agent ≠ Reviewer Agent ≠ Test Agent ≠ Release Approver. No single agent chain writing, reviewing, testing, and approving its own code.
Skills Registry
Company skills in a dedicated Git repo:
company-ai-skills/
├── product/
│ ├── requirement-analysis/
│ └── acceptance-criteria/
├── architecture/
│ ├── architecture-design/
│ └── impact-analysis/
├── development/
│ ├── backend/
│ ├── frontend/
│ └── mobile/
├── testing/
│ ├── unit-test/
│ ├── integration-test/
│ └── e2e-test/
├── security/
├── release/
└── operations/Each skill: SKILL.md, scripts/, references/, assets/, agents/openai.yaml. Codex loads skills from admin-level directory (/etc/codex/skills) for company standards, and project-level (.agents/skills) for project specifics.
Skill Change Governance
Skills become "AI R&D policy". Change flow: Modify Skill → Submit MR → Skill Test → Evaluation → Owner Review → Merge → Tag Version → Publish. Versions like [email protected], [email protected]. Agent Run must record skill_version to answer later why quality differed across months.
MCP Gateway (IBM ContextForge)
All agents connect via a unified gateway, not directly to MCP servers:
Codex → ContextForge → Plane MCP, GitLab MCP, DB MCP, CI MCP, Test MCP, Argo MCP, Security MCP, OTel MCP, Knowledge MCPContextForge provides centralized discovery, governance, and observability.
Business-Semantic MCP, Not Raw Admin APIs
Never expose production_kubectl. Wrap as business actions:
deployment.get_status
deployment.prepare_release
deployment.start_canary
deployment.get_canary_metrics
deployment.request_promote
deployment.request_rollbackAgents execute only permitted business actions.
MCP Permission Model
Dimensions: User + Department + Project + Agent Type + Environment + Tool + Action.
Example BackendDeveloperAgent: Git.read/write Allow; DB.dev.read/write Allow; DB.prod.read/write Deny; Argo.deploy Deny.
OperationsAgent: Git.read Allow; Grafana/OTel/Argo.read Allow; Argo.canary Approval; DB.prod.readonly Allow; DB.prod.write Deny.
Unified Identity via Keycloak
All systems (Platform, Plane, GitLab, Grafana, ArgoCD) authenticate via Keycloak (OIDC/SSO). Departments: product, development, qa, operations, management, security.
CI as Deterministic Quality Arbiter
Codex can write code, tests, run tests, explain/fix failures. But Codex cannot decide its own pass/fail . Final gate: MR → GitLab CI → Build → Lint → TypeCheck → Unit Test → Integration Test → Contract Test → E2E → SAST → SCA → SBOM → Quality Gate. GitLab CI pipelines with self-managed runners are ideal.
Test Platform
Test Agent takes Requirement Baseline + Acceptance Criteria + Code Diff → Test Design Agent → Generate Test Plan → Playwright → Browser → Results. Store HTML Report, Screenshot, Video, Trace.zip, Console, Network. Playwright Trace Viewer preserves failed test evidence. Platform UI shows per-acceptance-criteria results across browsers with screenshots and trace links.
Security Gate
Pipeline: GitLab CI → SAST, Secret Scan, Dependency Scan, Container Scan, IaC Scan, SBOM → DefectDojo (centralized import, dedup, triage) + Dependency-Track (SBOM-centric supply chain risk). Gate rules: Critical > 0 → Block; High > threshold → Block or Security Approval; Medium → Allow + Technical Debt.
Release: No Direct kubectl from Codex
Correct flow: Codex Release Agent → Generate Deployment MR → Git → Release Approval → Merge → Argo CD → Canary → Production Verification. Argo CD (GitOps) keeps production state in Git and approval system, avoiding agent-held cluster admin rights.
Production Verification
Post-deploy: Deployment → Wait 5 min → Production Verification Agent → OTel/Grafana → Error Rate, Latency, CPU, Memory, DB, Business Metrics → Compare Baseline. Example: Pre-release payment success 99.82% → Post-release 96.21% → Gate FAIL → Temporal triggers Rollback Workflow.
Production Incident Closed Loop
Grafana Alert (Payment 500) → Incident Created → Operations Agent queries Logs, Metrics, Traces, Deployment → Finds Release 2.3.1 → MR !928 → TASK-327 → REQ-102. Platform shows incident with impact, linked release, requirement, agent run, probable cause, and actions: Start Diagnostic Agent, Execute Rollback, Create Fix Task. Loop: Incident → Bug → Development → Testing → Release.
Agent Self-Monitoring with Langfuse
Langfuse (self-hosted) records Agent Run, Token, Cost, Latency, Tools, Errors, Retries, Human Rejection, Final Result. Management dashboard: Monthly requirements 128, AI auto-dev 102, First-pass 82, Human returns 20, Test escapes 3, Production incidents 1. Per-agent success rate, human edit rate, token usage, cost.
Complete Requirement Sequence Diagram
End-to-end flow from Product creating requirement through Plane/Temporal, Requirement Agent (PRD + Acceptance Criteria), Product Approval (reject/approve loop), Architecture Agent, Dev+Ops Approval, Task Planner, Parallel Codex Agents (Backend/Frontend/Mobile) with worktrees/MRs, Reviewer Agent, GitLab CI, Security Gate, Test Agent (Playwright), QA Approval, Product Acceptance, Release Agent, Ops Approval, ArgoCD, Canary, Production Verify (PASS/FAIL → Release/Rollback), Operations, Feedback.
Platform Menu Design
Home
├── My Work
├── My Approvals
├── Agent Running
├── Risks
└── Project Overview
Project
├── Projects
├── Requirements
├── Versions
├── Milestones
└── Members
Product
├── Requirements
├── PRD
├── Prototypes
├── Acceptance Criteria
└── Requirement Changes
R&D
├── Technical Designs
├── ADR
├── Dev Tasks
├── Agent Runs
├── MRs
└── CI
Quality
├── Test Plans
├── Test Cases
├── Automated Tests
├── Evidence
├── Defects
└── Security
Release
├── Releases
├── Environments
├── Deployments
├── Database Changes
├── Canary
└── Rollback
Operations
├── Services
├── Metrics
├── Logs
├── Traces
├── Alerts
└── Incidents
AI Center
├── Agents
├── Skills
├── MCP
├── Runs
├── Evaluation
└── Cost
Admin
├── Users
├── Departments
├── Roles
├── Policies
├── Quality Gates
└── AuditRepository Structure
ai-rd-platform/
├── apps/
│ ├── web/
│ └── api/
├── services/
│ ├── agent-orchestrator/
│ ├── codex-executor/
│ └── workflow-worker/
├── packages/
│ ├── domain/
│ ├── plane-client/
│ ├── gitlab-client/
│ ├── temporal/
│ ├── mcp/
│ ├── evidence/
│ └── authorization/
├── workflows/
│ ├── delivery/
│ ├── release/
│ └── incident/
├── agents/
│ ├── product/
│ ├── architecture/
│ ├── developer/
│ ├── reviewer/
│ ├── testing/
│ └── operations/
├── infra/
│ ├── helm/
│ ├── argocd/
│ └── terraform/
└── docs/
company-ai-skills/ (separate repo)
Business code remains in each team's own repos.Most Important Architectural Principle
The core is not Codex, not MCP, but the traceability chain :
Requirement Version → Agent Task → Agent Run → Code Version → Test Evidence → Release → Production Metrics → IncidentMust establish unique associations end-to-end. Click any production service version (e.g., Payment Service v2.8.31) and the platform answers: Why changed? (REQ-281), Who approved requirement? (Zhang San), Who designed? (Architecture Agent Run #882), Who developed? (Codex Run #912, #913), What Skill? ([email protected]), What MCP? (Git, Plane, DB-Test), What changed? (MR !822), Tested? (Unit ✅, Integration ✅, E2E ✅), Security? (Critical 0, High 0), Who approved release? (QA Li Si, Product Wang Wu, Ops Zhao Liu), When released? (2026-08-08 22:32), Post-release health? (Error Rate normal, P95 normal, Business KPI normal).
Only then can the company adopt "AI-led development, human review as safety net" black-box R&D at scale. This system is not an AI coding tool; it is the company's AI R&D Operating System .
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
