Tagged articles

FlowServe

1 articles · Page 1 of 1
Architects' Tech Alliance
Architects' Tech Alliance
Sep 6, 2026 · Artificial Intelligence

xDeepServe on CloudMatrix384: Full-Stack Design for Large-Scale MoE Model Serving

This article details Huawei's xDeepServe system for deploying massive Mixture-of-Experts models on the CloudMatrix384 supernode, covering the XCCL communication library, FlowServe decentralized serving engine, Transformerless execution architecture with prefill-decode and MoE-Attention decoupling, and hierarchical fault tolerance, achieving 2400 tokens/s per chip at 50ms TPOT.

Ascend 910CCloudMatrix384EPLB
0 likes · 18 min read
xDeepServe on CloudMatrix384: Full-Stack Design for Large-Scale MoE Model Serving