Tagged articles

CloudMatrix384

2 articles · Page 1 of 1
Architects' Tech Alliance
Architects' Tech Alliance
Sep 6, 2026 · Artificial Intelligence

xDeepServe on CloudMatrix384: Full-Stack Design for Large-Scale MoE Model Serving

This article details Huawei's xDeepServe system for deploying massive Mixture-of-Experts models on the CloudMatrix384 supernode, covering the XCCL communication library, FlowServe decentralized serving engine, Transformerless execution architecture with prefill-decode and MoE-Attention decoupling, and hierarchical fault tolerance, achieving 2400 tokens/s per chip at 50ms TPOT.

Ascend 910CCloudMatrix384EPLB
0 likes · 18 min read
xDeepServe on CloudMatrix384: Full-Stack Design for Large-Scale MoE Model Serving
Architects' Tech Alliance
Architects' Tech Alliance
Sep 5, 2026 · Artificial Intelligence

CloudMatrix384: Co-Designing Supernodes for Trillion-Parameter LLM Inference

Huawei's CloudMatrix384 integrates 384 Ascend 910 NPUs with a unified bus network and decoupled PDC service architecture to achieve 4.45 tokens/s/TFLOPS prefill efficiency on DeepSeek-R1 671B, outperforming H100 baselines through fused MoE communication, INT8 quantization, and heterogeneous pipelining.

Ascend 910CloudMatrix384Expert Parallelism
0 likes · 20 min read
CloudMatrix384: Co-Designing Supernodes for Trillion-Parameter LLM Inference