Architects' Tech Alliance
Sep 6, 2026 · Artificial Intelligence
xDeepServe on CloudMatrix384: Full-Stack Design for Large-Scale MoE Model Serving
This article details Huawei's xDeepServe system for deploying massive Mixture-of-Experts models on the CloudMatrix384 supernode, covering the XCCL communication library, FlowServe decentralized serving engine, Transformerless execution architecture with prefill-decode and MoE-Attention decoupling, and hierarchical fault tolerance, achieving 2400 tokens/s per chip at 50ms TPOT.
Ascend 910CCloudMatrix384EPLB
0 likes · 18 min read
