AI Coding: Escaping the Blind Box — Contracts and Gates for Verifiable Frontend Development

The author argues that AI can generate 90% of complex calendar UIs but the remaining 10-20% becomes a blind-box guessing game because developers cannot verify fixes; they propose four workflow gates (contract, main path, link, container) and executable skills to replace random patching with provable correctness.

Frontend AI Walk
Frontend AI Walk
Frontend AI Walk
AI Coding: Escaping the Blind Box — Contracts and Gates for Verifiable Frontend Development

Introduction

After completing a calendar feature with view switching, lists, details, and create/edit flows, the author compared the AI's first draft with the final production version. The initial skeleton covered the main flows and looked nearly complete. However, the subsequent bug fixes consumed disproportionate time: 2px radius mismatches, WebView background lines leaking, navigation stack clearing state, overlay buttons persisting without fields, images lost on edit refill, and occasional bridge callback failures on the app side.

The core suspicion emerged: AI may reach 90% completion, but the last 10-20% involves "AI blind-box draws" because the developer cannot verify whether a given change is actually correct. The model always produces an answer — either from the prompt or its own interpretation — and those interpretation gaps become the root cause of later bugs.

Core Essence: Not "Can't Write" but "Can't Verify"

The calendar page is actually a state network: same data rendered at different densities in day/week/month views; selected date, current view, filters, and scroll anchors get flushed on back navigation, edit, or new container opens; detail/create/blessing/policy/customer selection form a cross-page chain, not single-page CRUD; the container is not a browser — native WebView, bridge, and dual-platform differences break "seemingly reasonable" frontend logic.

AI excels at the first segment: translating requirements, interaction specs, and UI mockups into a "looks like a calendar" page structure. The bottleneck is the second segment: you cannot arbitrarily change code and cheaply prove "this cut fixed it?"

A typical late-stage pattern appears:

AI draft covers main flow → you think "only a few bugs left" → reality: still far from "sign-off ready" long tail.

Continuous small fixes → you think "steadily approaching correct" → reality: some commits random-walk in trial space, causing chained regressions.

Another round of style tweaks → you think "just visual polish" → reality: visual issues often entangle with container/state issues, leading to wrong fixes.

Tell AI "fix per this phenomenon, this solution" → you think "model will locate root cause" → reality: model often gives "compiles and passes once" local patches that aren't comprehensive.

One sentence: the bottleneck is not code generation speed, but the supply speed of "AI correctness evidence." Without evidence, humans are forced to substitute AI blind-box draws for judgment. Fixing one issue often introduces others due to insufficient consideration, leading to more blind-box draws until the human finally reads the code and crafts a comprehensive fix.

Core Questions: These Are Not the Same Problem

1. Which Layer Is AI Stuck At?

At least four layers exist; lumping them as "AI fails" is meaningless:

Structure layer : day/week/month, cards, overlays, entry points — AI usually sufficient.

Rules layer : all-day events, lunar holidays, empty states, field visibility, button appearance conditions — needs contracts, otherwise AI guesses.

Link layer : edit return, new container, query restore, cross-page refill — AI tends toward "locally correct, globally crashes" because these requirements aren't connected end-to-end.

Container layer : iOS/Android/HarmonyOS WebView, bridge callbacks, double-tap blank, system gestures — undocumented knowledge, model defaults weak.

Using "structure layer completeness" to estimate "overall AI completeness" yields inflated 80-90% figures. Professional testers on real devices uncover bizarre issues stemming from AI's overlooked details or incomplete understanding. AI always gives an answer — whether from your instruction or its own inference — and those inference differences are the core driver of subsequent bugs.

2. Why Does Late-Stage Fixing Become AI Blind-Box?

Testers and product managers feed back phenomenon sentences , not causal sentences , let alone answers :

"Page re-initializes after back navigation"

"This black line still appears on this section"

"Image not passed, not initialized"

"Button click no response / wrong navigation"

"Page errored"

"Why did it jump to login page"

Phenomenon sentences can drive AI to immediately write code, but they are insufficient for the human to decide: should I fix routing, lifecycle, native open mode, or style stacking? When you cannot pick a "unique correct solution," you let AI write a patch — essentially asking AI to sample. Sometimes the sample hits and fixes the issue; sometimes it introduces the next bug ticket.

3. Add More Skills or Fix Workflow First?

The author's verdict: Skills solve "less chatter on same-class problems next time"; workflow solves "can we verify this time?" Most pain in this calendar project's long tail came from the latter. Even beautifully written skills won't stop blind-boxing without "reproducible before change / comparable after change" AI constraint gates. The author acknowledges this is hard and may persist as long as AI can't solve 100% of complex needs; the difference is only how many blind-box draws you endure.

Core Elements: Four Categories That Actually Consume Time

1. Multi-View State Machine, Not Three Tabs

Day/week/month are not reskins. They share selected date, filters, data source, but interaction density differs; layered with "enter at specific timestamp", "restore on back", "fall back to original position after edit", state drifts from component memory to URL to container navigation stack. AI's initial design often keeps state in page memory — works with mock data, but breaks when real API responses arrive. Element: authoritative state source must be explicitly agreed (memory / query / native params) and "what to do when flushed" written down. Development docs must spell out conventions; cannot let AI guess.

2. Cross-Page Links More Fragile Than Main Page

The core calendar page is just the hub. The real fragmentation: create/edit refill, detail navigation, conditional buttons, empty field hiding, exception prompts, unified tracking timing. These changes have tiny diffs but depend on "what previous page passed, when this page requests, container push vs replace". AI excels at "fill what's missing on this page"; not at automatically maintaining "invariants across the whole chain". Element: write acceptance by chain, not by page. Refine features, refine requirements, refine acceptance criteria — by functional flow, not single page.

3. Visual Acceptance ≠ CSS Tweaks

Many late commits look like style optimization: radii, duplicate borders, floating buttons, note/tag layout, month view text truncation. Partly AI's MCP reading UI mockup fidelity issues; partly container background, overflow clipping, dual-platform rendering differences producing "fake style bugs". If you tell AI "make radius match design" without giving "parent background / safe area / WebView base color" context, it will modify at the wrong level. Element: visual issues — first determine level (page / component / container), then adjust properties. UI mockup interpretation still not 100%; when mockup hierarchy is deep, AI often powerless.

4. Native Capabilities Are Hidden Dependencies

Text language detection, copy succeeds but no callback, launch external app, edit page doesn't return to expected page — these aren't calendar algorithm problems, they're "H5 living inside native App" problems. If AI writes await bridge() per web common sense, it deadlocks on "success but no callback" bridges. For these pits, documentation and historical experience outweigh model weights. Element: native capabilities must have "failure appearance" and "success but no callback" patterns, not just success paths. AI often has understanding gaps and solution gaps in native-H5 bridging and WebView; these lurk in the dark and make debugging painful when exceptions hit.

Future Improvements: Workflow First, Skills Second, Stop Worshipping Percentages

The author presents a decision matrix:

Workflow — reduces blind-box draws, improves "provability of fix"; does NOT make model suddenly understand your native pits.

Skills — turns verified pits and constraints into triggerable default actions; does NOT replace current acceptance evidence.

Longer prompts — occasionally adds context; does NOT solve systemic unverifiability.

Stronger model — raises structure/rules layer hit rate; does NOT auto-fill your container layer's tacit knowledge.

Workflow solves "can we verify"; skills solve "fewer repeats of same pits". Don't reverse priority. Only after clarifying can you know how to improve; otherwise it's headless fly collision.

Improve Workflow: Add Four Gates for the Long Tail

Not more docs, but more "pause points":

Contract Gate (before coding) : Write day/week/month state, return restore, all-day/normal events, empty states, button visibility conditions as checkable clauses. No clause? No direct "let AI tweak".

Main Path Gate (during integration) : Only verify: visible, clickable, storable, returnable. Freeze main path first, then open visual and edge cases.

Link Gate (pre-test) : Record shortest reproduction script for "create → detail → edit → back → re-enter". Every AI patch must declare which chain it affects. Must consider full flow for acceptance.

Container Gate (final acceptance) : Run iOS/Android/HarmonyOS each once for bridge-related cases. If style issue appears only on one end, suspect container first, then CSS.

Without these four gates, you're using "commit count" to substitute "AI confidence".

Improve Skills: Write "Executable Pits", Not "Inspirational Norms"

Worth distilling into skills — things already paid for:

Calendar state sync to query trigger timing and write-back forbidden zones.

When edit scene returns to previous page vs opens new WebView.

Native bridge "success but no callback" coding pattern.

Overlay button and field null visibility rule templates.

Visual acceptance level checklist (container first, then component).

Not worth skilling first:

"Write high-quality maintainable code" platitudes.

Pasting entire UI spec into skill without trigger conditions and failure appearances.

Skill validation is simple: next same-class error, do you ask one less round, draw one less AI card? The author wants to stop repeatedly fixing persistent anomalies, trusting AI to handle recurring problem classes.

On the "90%" Number

Stop using "AI completed X%" as a management metric for model capability, workflow capability, or skill writing capability. More honest breakdown:

Demo completeness : can main path be explained clearly (often not low).

Acceptance completeness : can clauses be checked off (usually significantly lower).

Maintainability completeness : would next person dare touch state machine (often lowest).

Your gut feeling "maybe not 90%" mostly measures acceptance completeness ; the 90% spoken at demos is mostly demo completeness . They're not the same thing. A 100% number wouldn't mean zero issues. Better measure: how many fewer AI blind-box draws.

AI blind-boxing is a double-edged sword: feels good, but when pain comes it still delivers pain. The model can keep getting stronger, but the human directing AI must still answer: "Why do you think this AI change isn't another blind-box draw?"

Mapping to ai-frontend-dev-workflow: Fix Workflow or Skills?

Demo-able ≠ Acceptable. Workflow already good at turning requirements into first version; what's truly missing is how to keep using gates to constrain changes after integration, and whether cross-page links, container/bridge evidence can enter the acceptance denominator.

The author used the workflow's "mini mode", not "full mode", misjudging the project as simple — details turned out overwhelming.

Problems

Long tail falls out of workflow : marked "pending integration" as delivery closure, later bug fixes become chat blind-box draws.

Verify page, not chain : single page passes, but back/refill/re-enter whole chain crashes.

State and container lack contracts : view dimension, selected date, WebView return strategy left to guessing.

Phenomenon sentences drive code directly : "re-initialized again" → immediate patch, no root cause attribution.

Skills untestable, fixes un-attributable : interaction reports pass gates but key controls not exercised; Maker moves multiple chains at once.

Solutions

Fix workflow gates first, then harden relevant skills. Reversed order only accelerates AI blind writing. Constraints and gates are paramount; free-range ideas carry a price — the price of AI's freedom is higher.

Workflow changes : pending-integration checklist recallable, chain acceptance, authoritative state source into pre-coding gate, container gate, phenomenon sentences first clause-ified.

Skill changes : interaction analysis (control-level real test), contract open questions, Maker (single chain/level/bridge), 4b AC types (STATE/NAV/BRIDGE), pack by chain, experience preload.

These have been landed into the team marketplace's ai-frontend-dev-workflow (from IMP-073); field-level rules not expanded here, only the judgment retained. The changes aim to reduce future mud-wading when using the workflow, though the author cannot guarantee zero repeat mud-wading — may need several more passes to solve this class of problems.

Simple Summary

After using the workflow and encountering a series of problems, we only need to do three things:

Translate "phenomenon sentences" into "clause sentences" before feeding AI. From "page re-initializes after back" to "back must restore view + date; forbid clearing on remount". If prompt can't clarify, don't change code yet.

Each patch allowed to touch only one chain. Simultaneously modifying style, routing, bridge loses attribution ability, accelerates AI blind-boxing.

Write verified container pits and state pits as short skills; write acceptance gates into workflow. Gates first, then skills. Reverse order makes skills only accelerate blind writing.

These three things matter. They solve a series of downstream problems from AI programming.

This project didn't prove "AI can't write complex logic-heavy pages." Back to the opening scene: AI first version and production version side by side, skeletons nearly identical, gap entirely in the long tail. It more sharply proves another thing: when page is state network + link network + container network stacked together, generation is no longer scarce; scarce is "how do I know this version is correct?"

Structure layer can quickly look like 90%; acceptance completeness stalls at rules, return, refill, bridge — places where "fixed now but can't verify now". When evidence is insufficient, humans use blind-boxing to pretend progress — not a personal capability issue, but a method gap.

So the real takeaway: never think "next time switch stronger model" — at that moment the model isn't the key. When evidence for the problem is lacking, developers use AI blind-boxing to fix and push requirements. Not a personal capability issue, just a problem-solving method gap. Because I don't know what else to do, I must AI blind-box. Because time's up, can't not blind-box!

The model can keep getting stronger, but the person letting AI solve problems must still answer: "Why do you think this AI change isn't a blind-box draw?" The author hopes to reduce such development experiences. Fixing too many bugs and missing features makes you doubt yourself, doubt model capability, doubt workflow capability, doubt skill issues. We should believe in AI's capability, but we must learn to make AI make fewer mistakes. Let's keep pushing together!

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

state managementAI-assisted developmentfrontend workflowskill distillationverifiabilitycontract-driven developmentblind-box debuggingWebView integration
Frontend AI Walk
Written by

Frontend AI Walk

Looking for a one‑stop platform that deeply merges frontend development with AI? This community focuses on intelligent frontend tech, offering cutting‑edge insights, practical implementation experience, toolchain innovations, and rich content to help developers quickly break through in the AI‑driven frontend era.

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.