CloudMatrix384: Co-Designing Supernodes for Trillion-Parameter LLM Inference
Huawei's CloudMatrix384 integrates 384 Ascend 910 NPUs with a unified bus network and decoupled PDC service architecture to achieve 4.45 tokens/s/TFLOPS prefill efficiency on DeepSeek-R1 671B, outperforming H100 baselines through fused MoE communication, INT8 quantization, and heterogeneous pipelining.
