Taobao Overseas Tech's AI-Native R&D Shift: SDD, Agents & 85% AI Coding

Taobao and Tmall's overseas technology team shares their transition to an AI-native R&D paradigm using Specification-Driven Development (SDD) and Agent-driven collaboration via DingTalk groups, achieving 85% AI coding adoption and reducing end-to-end delivery cycles by 20-30% for medium-large requirements while maintaining quality baselines.

AliExpress Tech
AliExpress Tech
AliExpress Tech
Taobao Overseas Tech's AI-Native R&D Shift: SDD, Agents & 85% AI Coding

AI Coding Evolution and SDD-Driven Initial Results

AI coding has evolved from line-level IDE completions to conversational assistants and now to Agent-driven autonomous coding. The team adopted Specification-Driven Development (SDD): write design specs in natural language, then let AI generate code from specs and the existing codebase. Specs become the "instruction"; code is the artifact, not the starting point.

Quality is ensured through a phased gating process with human reviews at six stages: ① Project Preparation → ② Requirement Clarification → ③ Technical Proposal Generation → ④ Proposal Execution → ⑤ Test Execution → ⑥ Closure & Archiving. AI coding share rose from 5% to 85%, and overall delivery efficiency continues to improve.

SDD initial results dashboard
SDD initial results dashboard

The Perception Gap: Individual Speed vs. Organizational Throughput

Despite 85% AI-generated code for both frontend and backend, business and product stakeholders still feel delivery hasn't accelerated. Data shows coding occupies only ~20% of the end-to-end lifecycle; a 200% coding speedup yields only ~10% overall gain. Most time is spent in requirement alignment, PRD writing/review, technical design/review, integration testing, and cross-team coordination. The real bottleneck lies in the serial "Business → Product → Tech" handoff chain.

Lifecycle breakdown showing coding at 20%
Lifecycle breakdown showing coding at 20%

Agent-Driven Paradigm: Unified Interface, AI Execution, Human Decision

Core logic: align business/product/tech work surfaces, reduce communication overhead. DingTalk (ChatBox) is the proven collaboration paradigm; embedding an Agent brain turns it into an end-to-end execution platform. The traditional serial flow (Business requests → Product designs → Tech delivers) is reconstructed into "AI executes + Human decides" within a single DingTalk group.

Overall solution architecture
Overall solution architecture

Overall Solution

One-sentence trigger, avatar auto-relay: Operations states a business direction in natural language in the group; four AI avatars (Product, Tech PM, R&D, Analyst) automatically pick up and run the full chain from clarification to coding to data analysis.

Humans only at key gates: Three human checkpoints — PRD confirmation, AI test acceptance, final human sign-off. Avatars "get things done"; humans "judge correctness and releasability".

Clear engineering path: Every stage has a defined Agent carrier, dedicated knowledge base, and required engineering capabilities — mature components, not a demo.

Data feedback loop: Post-launch business data flows back; Analyst avatar generates strategy suggestions, triggering the next iteration.

Engineering guardrails throughout: Collaboration bus, artifact traceability, cross-review, mandatory gates for high-risk actions ensure controllability, trust, safety, evolvability.

Core Process

DingTalk group as hub; AI executes, humans decide. Agent drives requirement → design → clarification → delivery in a closed loop. All documents, meeting notes, and decisions stay in the group, forming a living domain knowledge base that continuously improves delivery speed and quality.

Three key implementation paths: ① Unified work interface ② AI-driven process collaboration ③ Automated delivery verification.

Core process diagram
Core process diagram

Key Paths

Channel: DingTalk group is the single work surface and context carrier; all docs, emails, meetings settle directly in the group.

Roles: AI produces (drafts, retrieves, summarizes, codes); humans decide (clarify, trade-off, approve, canary).

Flow: Requirement → Design → Clarification/Review GATE → Delivery → Release GATE — all executed inside the project group.

Targets

Focused on end-to-end delivery efficiency, with result and process metrics.

Result metrics (by complexity):

Faster: Small needs (≤5 person-days) zero R&D investment, Agent end-to-end same-day release; medium/large needs 20-30% cycle compression vs. baseline.

More: Same headcount, 40% monthly delivery throughput increase, validating capacity release from zero-R&D small needs.

Quality floor: Bugs per KLOC no worse than current; online incident and rollback rates no worse than baseline.

Process metrics (paradigm adoption health):

Paradigm penetration: % of needs delivered via new paradigm; Phase 1 target 100% for small needs in merchant supply domain, 50% for complex, then expand.

AI draft adoption: MRD/PRD final version evolved from AI draft (not rewritten), target ≥80%.

Clarification/review efficiency: Auto-instrumented group telemetry (rounds, gate durations) to baseline and reduce; MRD/PRD currently manual, will measure after AI assist.

AI coding share: Small needs from ~90% to 99-100% (near full automation); medium/large from 60% to 90%.

Key Capabilities

2.2.1 Unified Work Interface

DingTalk group integrates plugins, knowledge bases, databases. One project = one group; created at kickoff. Business, PD, PM, R&D, QA, and AI Agents all join. Group messages, files, cards auto-form project memory. Agents participate as group bots on equal footing.

Unified work interface
Unified work interface

2.2.2 AI-Driven Process Collaboration

AI drives, humans decide; process pulled back into the group. Business, PD, R&D collaborate in parallel. AI Agents clarify requirements, generate MRD/PRD drafts and technical designs, auto-archive discussions and decisions. Rhythm shifts from low-frequency big reviews (meeting → wait for doc → meeting) to high-frequency micro-cycles (AI draft → human quick review → AI immediate iterate). For complex coding, developers still prefer local IDE; the process guides them back to IDE for vibe coding, then return to group for updates.

Two tracks by complexity:

≤5 person-days: coding, testing, acceptance fully in-group; AI produces Diff and Test Report, aiming for one-shot success.

>5 person-days: group focuses on MRD/PRD/design review, upstream/downstream sync, progress tracking; vibe coding done locally in IDE, then results posted back.

Unified principle: regardless of track, requirement progression, artifacts, and key decisions stay in-group, flowing as card messages with instant traceability.

AI-driven collaboration flow
AI-driven collaboration flow

2.2.3 Automated Delivery Verification

Under SDD, PRD and technical design merge into machine-executable Spec. AI runs coding, unit tests, regression, deployment, verification, release automatically. Humans guard only three gates: design approval, core code review, canary release approval.

Two tracks again:

≤5 person-days: full closed-loop in DingTalk, no context switch, no queue.

>5 person-days: DingTalk as process cockpit; vibe coding and deep testing/regression back in local IDE — keeps automation speed while retaining control for complex engineering.

Every pipeline stage hangs on objective quality gates — unit test, lint, security, regression, online metrics — anomalies auto-rollback. Human gates are risk-graded from L1 (lightweight) to L4 (full human decision). AI only computes impact scope; decision authority returns to humans.

Automated delivery verification pipeline
Automated delivery verification pipeline

Practice Path and Results

2.3.1 Practice Path

Pilot scenario selection: Start with "what-you-see-is-what-you-get" simple needs — copy/button tweaks, sorting rule changes, config updates. Simple PRD, enumerable design, low acceptance bar; easiest to validate "one sentence → release" in-group loop. Then expand to medium-complexity features.

In-group gate mechanisms: Standardize "permission check, risk identification, output quality" as AI gates in-group: every Yes/No must verify approver identity; AI self-checks drafts (completeness, conflicts, missing deps) before human review; high-risk changes force human decision.

Measurement & feedback loop: Every group flow instrumented — stage latency, AI draft adoption, human rework count, final quality. Data feeds back into Skill prompts, process templates, forming "real business runs data → data improves AI → capability expands pilot scope" virtuous cycle; also seeds digital avatar datasets.

2.3.2 Phase Results

New paradigm validated in Supply, Payment, and Marketing domains. Capability extended from code assistance to requirement analysis, design generation, R&D execution, project coordination — proving end-to-end delivery and cross-role collaboration feasibility.

R&D Delivery: Supply domain small needs achieved full in-group delivery from MRD/PRD/design through coding, deployment, test, release. AI corrected technical design via code analysis, cutting overall R&D estimation by 50%. Complex needs (e.g., Panama) use hybrid "group collaboration + IDE development" mode, supporting multi-module CR, test case generation, pipeline tracking, pre-prod issue triage.

Project Coordination: Payment and Supply domains connected meeting minutes, PRD clarification, todo tracking, technical design, R&D delivery. Meeting conclusions auto-convert to action items with owners and deadlines, flowing into design, CR, code delivery — continuous sync of progress, risks, todos, reducing scattered info and repeated alignment.

Requirement Analysis: Marketing domain completed current-state mapping, feasibility judgment, upstream/downstream impact analysis on real needs, generating PRD drafts and interaction demos. Product, business, R&D discuss on same analysis artifact, cutting early-stage back-and-forth.

Future Iterations

Process End-state: Operations initiates in group; multi-role AI avatars auto-relay through clarification, design, dev, release, data analysis; humans only at final acceptance; business enters continuous self-iterating loop.

Data Flywheel: Post-launch data auto-feeds back, calibrates knowledge bases, optimizes avatar behavior, evolving system from "one-off delivery" to "continuous self-iteration and refinement".

Management & Intervention: Build unified admin console for monitoring, rule definition, strategy analysis, manual intervention — ensuring "observable, controllable, correctable".

Digital Employees: Beyond per-need "avatars", create "digital employees" that reside permanently in a business domain with defined responsibilities and authority boundaries. They proactively claim eligible needs, complete within boundaries, escalate only when out of scope. Avatars own a single delivery; digital employees own continuous delivery for a domain.

Future iteration roadmap
Future iteration roadmap

Summary: L1–L3 Maturity Levels

AI-era R&D paradigm classified by AI participation into L1–L3. After one month of real delivery practice, Taobao Overseas Tech achieved:

Full-cycle delivery validation; business/product/tech aligned on collaboration flow.

Communication and collaboration efficiency uplift across roles; eliminated fragmentation-induced information gaps.

R&D paradigm upgraded: ≤5 person-day needs fully delivered in-group (coding, deploy, test).

Small needs essentially at L3; most needs still at L2. Next focus: extend L3 to medium/large needs via four workstreams:

Process Specification: SDD spec as input, fused with AAIC Skill hosting/scheduling to drive "coding & integration & test", gradually automating parts that currently require local IDE; humans keep only risk-graded gates.

Document Quality: Iterate prompts, example libraries, project/domain knowledge bases so MRD/PRD/design AI drafts reach review-ready quality, reducing rewrites.

Engineering Foundation: Polish Skill orchestration, bot latency, group card interactions; complete identity verification for every Yes/No and cross-org authorization.

Asset Iteration: Telemetry on stage latency, AI draft adoption, human rework; group info沉淀 as digital assets — traceable, replayable, reusable — continuously enriching and calibrating knowledge bases.

When these four flywheels spin, the L2/L3 boundary truly sits on scheduling authority: process engine owns stage transitions; AI autonomy expands from single-stage to cross-stage end-to-end; humans arbitrate only when risk materializes. Metrics shift from AI code share to end-to-end cycle time — throughput no longer linearly tied to headcount, but to knowledge base quality, domain Skill coverage, and trustworthiness of permission/boundary controls.

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

AI codingR&D efficiencyDingTalkSDDAgent-driven developmentend-to-end deliveryTaobao Overseas Tech
AliExpress Tech
Written by

AliExpress Tech

Official tech channel of AliExpress International Tech Division, showcasing the latest technology developments and innovations in global e‑commerce.

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.