R&D Management 28 min read

Why a Working Prototype Isn't Production Ready: FDE's First Controlled Loop

This article explains how Forward Deployed Engineering (FDE) bridges the gap between prototype and production by defining a controlled, runnable, verifiable loop, classifying prototypes by validation goals, establishing iteration gates, and requiring explicit keep/discard/block lists before scaling.

Data Bricklaying Diary
Data Bricklaying Diary
Data Bricklaying Diary
Why a Working Prototype Isn't Production Ready: FDE's First Controlled Loop

Introduction

Previous articles in this series established an evidence pack from field discovery, converged field knowledge into a business source model, and clarified system boundaries, authorities, interfaces, and fallbacks. The team can now explain "why this design." The next hard question: what should the first runnable thing look like?

Many projects lose boundaries at this step. A prototype demo succeeds, then the team adds features, connects real data, and even requests production permissions. Hardcoded samples, temporary scripts, and one-off accounts remain in the system. Only when someone asks about monitoring, failure takeover, rollback, and ownership does the project realize it has merely grown from a demo into a larger demo.

This article continues using the "manufacturing enterprise abnormal order" synthetic scenario. In the boundaries set by Article 04, the first phase only reads controlled order and inventory states, identifies potentially delayed orders, generates risk tasks with evidence, and requires authorized personnel to confirm. It does not automatically reschedule or change customer commitments.

FDE, after architecture boundaries are fixed, must not enlarge the prototype but cut the first controlled, runnable, verifiable loop. It can enter real paths to collect feedback, but before the full production gate passes, it cannot claim production readiness or stable operation.

Prototype Is Not a "Small Production System"

Teams often treat a prototype as a juvenile production system: build main features fast, then gradually add permissions, tests, monitoring, and ops. The flaw is that prototypes and production systems answer different questions.

A prototype primarily reduces uncertainty. It may use limited samples, isolated environments, manual operations, and temporary UIs to quickly judge whether a direction is worth pursuing. A production system must face real permissions, abnormal inputs, capacity fluctuations, dependency failures, version changes, and long-term responsibility.

Therefore, moving from prototype to production is not "keep adding features" nor "one engineering hardening before launch." It is repeated re-confirmation of scope, risk, and responsibility. What is allowed in a prototype may not enter the next phase; unverified assumptions do not become true just because a demo succeeded.

Earlier articles "How to Define the First AI Pilot" and "AI Pilot Cannot Stop at PoC" discussed scenario selection and full production admission criteria. This article answers the missing middle step: how FDE converges a prototype into a delivery slice that continuously gains real feedback without crossing the production declaration boundary.

Four Prototype Types Validate Different Hypotheses

"Prototype works" does not reveal what was actually validated. The author proposes classifying prototypes by their primary hypothesis:

Business Adoption — validates whether business personnel are willing to view, judge, or use results in real work nodes. Cannot prove data stability or AI capability.

Data Feasibility — validates whether required data can be obtained per agreed semantics, timeliness, and permissions. Cannot prove user adoption or business improvement.

AI Capability — validates whether the model or rules can complete the target judgment on representative data. Cannot prove system integration, permissions, operation, and reconciliation closure.

System Integration — validates whether interfaces, events, identity, errors, and human takeover can run through. Cannot prove AI quality and business value are established.

In the abnormal order scenario, each prototype type is designed accordingly:

Business adoption prototype only verifies whether planners are willing to review risk tasks and evidence in daily handling.

Data feasibility prototype only verifies whether orders, inventory, and commitment bases can be obtained per business semantics and timeliness.

AI capability prototype only verifies whether rules or models can identify delay risk in normal, abnormal, and conflicting samples.

System integration prototype only verifies whether order changes trigger tasks, confirmations are recorded, and failures can retry or fall back to human.

A prototype may incidentally expose other issues, but it must have one primary validation goal. Otherwise, a few prepared data points can simultaneously "prove" business acceptance, data availability, model accuracy, and system stability — the demo looks complete, yet no hypothesis gains verifiable evidence.

Before launching a prototype, the team must explicitly write: the hypothesis to validate, data scope, success and failure samples, validation environment, decision makers, and the continue/shrink/stop decision after the demo.

What Is the First Controlled Runnable Loop

This "controlled runnable loop" is the author's concept for this series, not an industry standard. It sits between prototype and production readiness, requiring three properties simultaneously:

Controlled — scope has not secretly expanded. Business objects, users, data, permissions, actions, and observation windows have explicit boundaries. High-risk writes retain human confirmation; unapproved data cannot be temporarily accessed; FDE cannot bypass customer identity and authorization with demo accounts.

Runnable — it can repeat execution in a controlled real path or high-fidelity simulation. "Real path" does not equal formal production. It can be a pilot environment with limited users/data, a shadow path, authorized data replay, or a read-only process without high-risk side effects. The key is that input, processing, result, human confirmation, and feedback no longer rely on demo personnel ad-hoc stitching.

Verifiable — every step leaves evidence. The team can prove which version was used, what data entered the path, what judgments were produced, who confirmed, where failures occurred, how recovery happened, and which backlog or next version received the feedback. One smooth demo is not verification; repeatable runs that identify failures and leave records are.

For the synthetic abnormal order case, the first loop compresses to five steps:

1. Controlled read of order and inventory state
2. -> Identify potentially delayed orders
3. -> Generate risk tasks with source and rule version
4. -> Authorized planner confirms, rejects, or supplements evidence
5. -> Record results, failures, and feedback

This loop intentionally excludes automatic rescheduling and customer commitments, and does not cover all order types. The scope is small not because the team cannot do more, but because new write permissions, irreversible actions, and cross-department responsibilities are not yet closed.

After the loop, the team only gains a limited declaration: this path can run controlled and be continuously verified under current scope and conditions. It does not represent that capacity, security, operations, governance, and long-term value have passed the full production gate.

Three Mandatory Actions When a Prototype Ends

The most dangerous moment for a prototype is not failure, but when results look good and everyone wants to "just continue on this basis." FDE must drive an explicit closure.

First, Decide What to Keep

Keep not all written code, but capabilities with evidence support that can enter the loop and their verification assets.

In the abnormal order scenario, keep: controlled read of orders and inventory, delay risk judgment, evidence references, risk tasks, human confirmation records, and corresponding normal, conflict, and failure samples. Each capability must be traceable to the business model, architecture decisions, interface contracts, and tests.

Second, Clarify What to Discard

Hardcoded orders for the demo, one-off datasets copied to personal machines, temporary scripts, shared accounts, manually modified return results, and UIs serving only the demo path must not naturally flow into the next phase.

"Discard" does not mean delete all prototype code, but explicitly mark which assets lack inheritance qualification. Implementations with reference value can be kept in isolation but must not continue to be called nor packaged as formal capabilities.

Third, Write What Blocks Productionization

Blockers must be actionable or escalatable facts, not "continue optimizing." Examples:

Auto write-back of commitments not approved by business and permission owners.

Failures unmonitored, task loss unreconciled.

Irreversible actions lack human takeover and compensation paths.

Key dependencies have no owner; interface changes unattended.

Temporary implementations have no expiry but are already continuously depended on by real users.

Closure deliverables must include a keep list, a discard list, and a risk blocker list. Without these three boundaries, the prototype never truly ends — it only becomes a harder-to-remove temporary system.

FDE Delivery Rhythm: Not Just Faster

Traditional product development organizes rhythm around roadmaps, version plans, and unified user groups. FDE sites constantly receive new business facts, data limits, permission requirements, and organizational constraints. Feedback is more direct; changes penetrate code and interfaces more easily.

Thus, FDE rhythm should not be "customer requests daily, engineers change daily." A better approach: shorter feedback cycles plus harder quality gates.

Each iteration must freeze five items:

Iteration boundary: which hypothesis is validated, which objects and users are covered, what is explicitly not done.

Quality gates: what contract tests, error samples, reconciliation, and rollback evidence are required to enter and exit the iteration.

Customer review cadence: what items are daily syncs, what items are decided periodically by business, system, data, or security owners.

Change management: when site feedback changes requirements, model, or system boundaries, it goes to the backlog, project decision log, or ADR.

Temporary implementation convergence: who owns temporary scripts, configs, and manual steps; when they expire; whether they enter formal contracts or are deleted.

The "Forward Deployed Engineering Practice Guide" defines delivery rhythm as ensuring key facts reach the right responsible persons at the right time, distinguishing daily pilot checks, value/risk reviews, expansion reviews, and release readiness reviews. This is the guide's suggestion, not an industry standard. The article adopts its core idea: rhythm must serve decisions; every review must produce explicit evidence and responsibility actions.

Using the abnormal order loop as an example: daily checks focus on task loss, error judgments, human rejections, and data anomalies; phase reviews decide continue verification, shrink order scope, adjust rules, or stop investment. Changes involving system boundaries, data authority, or interface contracts go to ADR; only iteration scope and rhythm decisions go to the project decision log. Neither can be replaced by chat logs.

A minimal delivery rhythm table:

Daily Loop Check — Must answer: yesterday, any task loss, error judgments, human rejections, data anomalies? Output: feedback queue, failure samples, blocker owners. Escalation triggers: high-risk misjudgment, sensitive data exposure, core path unavailable.

Iteration Review — Must answer: was this round's hypothesis validated? Which capabilities kept, discarded, blocked? Output: hypothesis evidence, three lists, project decision log. Escalation triggers: insufficient evidence, rollback not rehearsed, temporary implementations unowned.

Expansion or Shrink Review — Must answer: expand users, order scope, or data scope? Output: expansion boundaries, observation windows, rollback paths, confirmation records. Escalation triggers: new scope unclear, human takeover capacity insufficient, authorization changes.

This table is not a new project management template; it binds feedback, evidence, and responsibility into the same rhythm. Every beat must have explicit participants and decision makers; lacking necessary evidence, the default action is pause expansion or shrink scope — not cover gaps with the next demo.

How FDE Site Rhythm Differs from Product Development Rhythm

Comparison across five dimensions:

Primary change source: Product — roadmap, market feedback, version planning. FDE — customer site facts, data, permissions, organizational constraints.

Primary review objects: Product — product, dev, test, ops teams. FDE — business owners, system owners, data/security roles, frontline users.

Temporary implementations: Product — usually avoid. FDE — may be forced by site constraints but must be explicitly managed.

Definition of done: Product — features, quality, release, product metrics meet version requirements. FDE — current slice controlled, verifiable, rollbackable, responsibility and next step clear.

Change records: Product — product backlog, requirements, version records. FDE — simultaneously maintain site evidence, project decision log, ADR, and customer confirmations.

Neither rhythm is superior; they are not replacements. FDE must respect the product base's version and compatibility responsibilities; product teams cannot use a unified roadmap to flatten customer site authority, permission, and operation differences.

Two losses of control must be prevented: (1) site feedback directly modifies product core, creating unmaintainable customer specials; (2) product rhythm ignores site blockers, forcing long-term reliance on manual patches.

Iteration Gates: When to Continue, Shrink, or Stop

Beijing policy references FDE, "on-site co-creation," and "continuous iteration feeding back intelligent agent capability improvement." This supports that one-time delivery is insufficient for intelligent agent landing, but "continuous iteration" does not equal infinite investment, nor can it replace specific project verification gates.

At each iteration end, the team must distinguish two signal types.

Signals to Proceed to Next Round

This round's primary hypothesis has verifiable evidence.

New risks accepted by authorized body or eliminated by scope reduction.

Failures detectable; rollback or human takeover rehearsed.

Customer willing to continue use in agreed real process.

Temporary implementations have owner, expiry date, clear destination.

Signals That Must Shrink or Stop

Data usage, identity, permissions, or high-risk action authorization not closed.

Failures unobservable; transaction results unreconciled.

Irreversible actions lack takeover and compensation paths.

Temporary implementations unowned but already in continuous use.

Conclusions still only from demo samples, not repeatable in controlled path.

Same key hypothesis unsupported for multiple consecutive rounds with no new evidence path.

Gates are not to add meetings but to turn "can we continue investing" into checkable facts. If not passed, default action is not "do another round and see" but shrink objects/users/data/actions, downgrade to human assist, fill evidence gaps, or stop investment and retain retrospective assets. Continue, shrink, and stop are all professional conclusions. Real failure is no evidence and no decision, letting a temporary system run indefinitely.

FDE Responsibilities in the Runnable Loop Phase

FDE is not a bystander collecting feedback beside the prototype. It must personally participate in key implementations, transform business models and architecture decisions into runnable slices, establish verification and failure records, organize authorized role reviews, and maintain tracking from temporary implementations to formal contracts.

Specifically, FDE must drive at least four things:

Every prototype has a clear validation goal; demos do not replace evidence.

Every iteration has scope, gates, review deliverables, and decision records.

Prototype assets are classified into keep, discard, block — none enter the next phase with defects.

When authorization, risk, or production gates are unclosed, actively shrink or stop — do not bypass.

This responsibility has a ceiling. FDE cannot decide business priorities for business owners, cannot approve permissions for data/security/system authorities, cannot write customer trials as proven production results, and cannot self-declare the loop production ready.

Loop Completion Signal Is Not "It Runs"

The first controlled runnable loop completes not because the happy path finally runs, nor because the customer approves after a demo.

Completion signals: validation goal explicit; loop repeatable and leaves evidence; failures detectable and takeoverable; kept capabilities, discarded demo assets, and production-blocking risks documented; every temporary implementation has an owner and destination; team can make continue/shrink/stop decisions based on gates.

At this point, the project still does not have a "production ready" conclusion — only a controlled path worth further verification.

Article 05 cuts this loop. The next article will address: when the loop includes enterprise Agents, models, knowledge, and tool calls, how is identity inherited, data authorized, actions controlled, evidence retained, and who takes over failures.

References

AI 项目第一个试点怎么定:从价值叙事到最小价值闭环 —

https://mp.weixin.qq.com/s?__biz=MzcwNDM2MTg3NA==&mid=2247484197&idx=1&sn=bd5410a22542e13de4b119d1fb392907&scene=21#wechat_redirect

AI 试点不能止于 PoC:为什么需要明确的“毕业条件” —

https://mp.weixin.qq.com/s?__biz=MzcwNDM2MTg3NA==&mid=2247484424&idx=1&sn=cfdfe550bef5c2c5b6d567621869aa19&scene=21#wechat_redirect

功能完成不等于生产就绪:AI 编程为什么更需要工程化 —

https://mp.weixin.qq.com/s?__biz=MzY4ODQxNzY2NQ==&mid=2247483726&idx=1&sn=c5ef2f9d8eb7a91dbdc4f24591049cb5&scene=21#wechat_redirect

别急着画架构图:FDE 如何从真实现场发现可工程化的问题 —

https://mp.weixin.qq.com/s?__biz=MzcwNDM2MTg3NA==&mid=2247484841&idx=1&sn=9ba55d22d8e8728c3dba9d1eaa7a4ab2&scene=21#wechat_redirect

业务专家不该直接面对技术标准:FDE 如何把现场知识变成可运行的业务模型 —

https://mp.weixin.qq.com/s?__biz=MzcwNDM2MTg3NA==&mid=2247485048&idx=1&sn=746d1c9b43f5bc6f532aa5b49f19b031&scene=21#wechat_redirect

架构图不是答案:FDE 如何在客户约束中做系统边界与集成取舍 —

https://mp.weixin.qq.com/s?__biz=MzcwNDM2MTg3NA==&mid=2247485149&idx=1&sn=e325cde65425680283592fea73d9e193&scene=21#wechat_redirect

前线部署工程实践指南:交付节奏与决策节拍 —

https://github.com/yeasy/forward_deployed_engineering_guide/blob/

前线部署工程实践指南:原型、试点与生产化 —

https://github.com/yeasy/forward_deployed_engineering_guide/blob/542fc68c225556a71cf162ea9ef21eae303ae7ca/04_delivery/4.3_lifecycle.md

北京市关于加快智能体引领发展的若干措施 —

https://www.beijing.gov.cn/zhengce/zhengcefagui/202607/t20260723_4781085.html
Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

risk managementFDEproduction readinessForward Deployed EngineeringControlled LoopDelivery RhythmIteration GatesPrototype Validation
Data Bricklaying Diary
Written by

Data Bricklaying Diary

Records practices, thoughts, and pitfalls on the data grunt-work journey, sharing content on data platforms, data analysis, data processing, data governance, knowledge graphs, and more. Less theory, more hands‑on, making complex data technologies simple.

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.