How Alipay Built the Agentic Terminal Engine Behind 'Abao'

Alipay's xUI technology stack powers the 'Abao' AI agent with streaming multimodal communication via MoQ, generative UI rendering across three generations (Markdown, HTML, A2UI), on-device agent execution via GUI/TUI/Workflow, and cross-agent collaboration through AHA protocol, enabling asynchronous, companion-style interactions across devices.

Alipay Experience Technology
Alipay Experience Technology
Alipay Experience Technology
How Alipay Built the Agentic Terminal Engine Behind 'Abao'

Mobile Agent Application Scenarios

Mobile apps are shifting toward Agent-based architectures. Over the past two years, Alipay has transformed service delivery in search, travel, and government scenarios using conversational Agents, and expanded ecosystem cooperation by embedding mini-programs into car infotainment systems and opening payment capabilities to external Agents. 'Abao' is positioned as the AI entry point for Agents — not merely a chatbot, but a combination of 'AI entry + massive business service supply + multi-end open ecosystem.' Supporting this requires solving four technical domains: continuous efficient communication between human and Agent, efficient friendly interaction and service completion atop communication, audio/video multimodal interaction, and on-device capability invocation and execution by Agents. These domains span client and cloud, and must also consider collaboration with other vendors' Agents.

Alipay's Agentic Terminal Practice

Before AI dialogue, client request patterns were one-off request-response. With AI, Agent content generation is continuous: a user utterance triggers a continuous stream of responses, requiring streaming communication. Users can change intent mid-conversation, demanding full-duplex links. When Agents perform tasks — screen operations, UI inspection, camera feed analysis — multimodal full-duplex streaming becomes essential for real-time perception of voice, video, and screen, followed by reasoning, tool calls, or user confirmation.

Streaming Multimodal Real-Time Transmission: Communicating with Agents

Early text dialogue emphasized efficient streaming of text content; simple data dialogues worked over gRPC duplex channels. Complexity grew with joint audio/video and digital humans. The team initially used RTC combined with other data channels, but RTC is designed for human-to-human interaction, requires room establishment with long connection times, and its extensibility centers on audio/video, making multimodal fusion and unified control inflexible.

The MoQ protocol offers a new approach designed for client-Agent communication. It multiplexes all modalities, natively supports interruption and other human-Agent coordination behaviors, and provides strong link, modality extension, and information flexibility. Alipay adopted MoQ as the foundation, complemented by RTC and other channels, preserving RTC's mature capabilities in specific scenarios while gaining flexible multimodal expansion via MoQ.

Generative Interactive UI: Interacting with Agents

The interaction expression layer evolved through three clear stages:

Stage 1 – Early AI Q&A: Models returned large text blocks. Markdown was popular because models generate it natively. The team built three-end native Markdown rendering (iOS, Android, Web) instead of embedding a browser engine, because native rendering better supports cross-platform standard protocols and guarantees AI interaction effects with lower engineering cost and higher stability.

Stage 2 – Complex Frontend Code Generation: As models grew stronger, they could generate complex frontend code. Dialogue content expanded beyond text cards to charts, analysis, knowledge visualizations, requiring browser rendering. The core challenges became performance, reliability, stability, and smooth animations. Alipay adopted its self-developed MYWeb kernel and deeply optimized both MYWeb and WKWebView, enabling Web blocks embedded in native dialogue UI while maintaining high experience quality.

Stage 3 – End-to-End Service Completion in Dialogue: The first two stages handled expression and simple exploration but could not complete transactions. Alipay introduced Google's A2UI (Agent-to-UI) paradigm: design interactions to complete a service, not generate pages. The model generates what to do , which is translated into UI expressions via standardized component composition, operation protocols, and information flows. Expression and Agent layers collaborate through protocols, decoupled. Under A2UI, the cloud Agent understands user intent, generates steps, translates them via MCP or Skill into staged UI cards rendered on-device, where users view, confirm, select, and complete the service.

This yields three generations of rendering specs: (1) Streaming Markdown for streaming effects and native rendering performance; (2) Streaming HTML rendering for complex expression, interaction, and stable engineering; (3) Declarative UI rendering combined with A2UI protocol and MCP supply, solving real task completion in dialogue. The stages progress from expression → complex expression → full transaction closure.

Audio/Video Call Technology: Interacting with Agents via Voice and Vision

Beyond text, users need voice and multimodal interaction. Stage 1: voice streamed to ASR, transcribed to text, sent to Agent, response TTS-played — a two-stage pipeline with poor experience. Stage 2: real-time answer — voice streamed, understood in real-time, fed to Agent for streaming inference, results streamed to TTS, forming full-duplex. Stage 3: richer multimodal expression — images, real-time video capture, real-time digital humans — fused into a multimodal real-time link with coordinated duplex continuity and interruption support. Combined with the network evolution above, this forms a complete multimodal interaction system.

The architecture splits into layers: top business layer, standard SDK for interaction control, generative rendering, and session management. For multimodal core, three sub-layers: (1) Media layer — state/session management, media capture/playback, hardware algorithms, production optimization; (2) Network abstraction layer — decouples media from network transport, isolating dependencies; (3) Transport layer — MoQ primary, RTC fallback, gRPC last resort, maintaining high transmission efficiency. This layering significantly improved connection latency, perceived usable latency, and network stutter rates. The network abstraction layer is key: media encoding/algorithms and network transport were deeply coupled; abstraction lets upper media capabilities evolve independently, while lower protocol swaps or degradation don't affect business experience. Example: weak networks fall back to gRPC for basic text; strong networks use MoQ for full multimodal. The standard SDK lets businesses onboard quickly without reimplementing complex session and media logic.

Terminal MCP, UI Agent Execution: Making Agents Get Things Done

A2UI works when real MCP and API supply exist (e.g., ride-hailing, coffee ordering). But not all services expose standard MCP. For those, Alipay uses GUI Agent-driven perception and execution: model's UI click intent converted to executable actions by fusing multiple on-device tech stacks — layout tree decomposition/merging, capture, click detection — simulating system-level click permissions inside the app without system capabilities.

Alternative: TUI (Textual UI) — page structure described as structured text (functions, positions) given to the model; model understands services via text, tells which structured location to click, no visual/absolute coordinates needed. GUI has higher generalization, continuously improvable via general GUI models. TUI has slightly lower generalization due to structured expression design, but faster execution and avoids complex visual reasoning overhead.

Third: model-free execution for fixed scenarios/flows (e.g., search fixed object, standard processes). Compose Workflow/scripts from the first two capabilities; when user intent matches, run directly. Fastest execution, lowest generalization, maintenance cost from business logic changes.

All execution modes need a standard management mechanism: on-device capability registry (standard GUI ops, per-scene Workflows/scripts discoverable by Agent), authentication/authorization to prevent rogue calls. Success rate is critical — GUI execution must exceed 90% to truly complete tasks. During end-cloud collaboration, UI may change anytime; the solution: prevent user clicks during execution, perform explicit verification before capture and execution — confirm page stability, confirm model's target still at original position — keeping error rates low. Without verification, models mis-click, context and operations become chaotic, experience degrades sharply.

Cloud side selects tools based on device states and calls, handles task rules and state management, schedules via MCP. Algorithm investment is heavy: off-the-shelf GUI models lack deep understanding of Alipay, government, and local mini-program layouts. Alipay iterates model understanding of massive mini-programs: define service descriptions, high-quality annotation, train at suitable scale, ensure inference speed for smooth per-step operation. Build algorithm innovation flywheel so data and scenarios continuously feed back into models for long-term generalization.

Overall strategy: fixed scenarios/flows → direct execution; complex scenarios → GUI; simple structured scenarios → TUI. Choose execution mode per scenario. Core is owning execution capability and continuous algorithmic model building.

AHA Agent Interconnection: Multi-Agent Collaboration

User intents aren't confined to Alipay. Distributed across apps, they need system Agent and app-side Agent collaboration. Example: user asks phone system assistant 'buy a movie ticket'; assistant understands intent, delegates to Alipay; Alipay Agent operates in background to purchase. User needn't watch screen — can browse other apps or chat elsewhere. On completion, assistant notifies user of final order page with async result. Trigger via text or voice; user can interrupt/cancel anytime. Demonstrates cross-app Agent collaboration experience.

In multi-Agent interconnection, intent distinguishes cooperation. Intent may arise in terminal system or inside Alipay; relationship is service distribution around intent. If intent originates in phone assistant, assistant finds corresponding Agent service, invokes it, hands off intent for domain-specific closure. This avoids vendor directly operating Alipay or Alipay breaching permissions to call system capabilities — clear responsibility split: system assistant understands intent; Alipay provides execution capability and service supply.

Vendor cooperation adds two benefits: (1) Alipay can complete tasks in background after assistant handoff, notifying user asynchronously; (2) system assistant's system-level capabilities are more standard and lower error rate than client-side simulation. Combined, they form a complete multi-Agent interconnection solution.

Agentic Terminal Development Outlook

Previous sections solved communication, interaction, execution, collaboration — delivering Alipay services via Agent. Further exploration:

Asynchronous execution: Knowledge queries and service triggers already async; but GUI-based service execution in Alipay still requires user to watch step-by-step. Goal: end-cloud collaboration to make these async too.

Companion-style interaction: Can users converse with Abao in any Alipay scene, with real-time scene understanding and perceptible assistance?

Multi-device extension: Beyond phones, Abao interacts with watches, glasses, payment terminals — extending service supply to these scenarios.

Cloud sandbox execution: Sandboxes running MCP, OpenClaw — can they run mini-program services directly? If yes, cloud can complete services for users, return results to device, supporting longer async tasks.

Companion experience focus: Shift from 'user finds service' to 'service finds user' — Abao everywhere. Requires continuous enhancement of async tasks, task collaboration, voice interaction on device; end-cloud multi-device, multi-scenario context management; cloud sandbox direct mini-program execution — together forming next-phase evolution path.

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

Workflow AutomationAgentic AIGUI AgentMulti-Agent CollaborationGenerative UITUIA2UIMoQ ProtocolAlipay xUIMultimodal Communication
Alipay Experience Technology
Written by

Alipay Experience Technology

Exploring ultimate user experience and best engineering practices

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.