Tagged articles

XCCL

2 articles · Page 1 of 1
Huawei Cloud Developer Alliance
Huawei Cloud Developer Alliance
Sep 8, 2026 · Artificial Intelligence

AReaL v1.0.5 LoRA RL: Low-Rank Adaptation for Accessible Large Model RL on Ascend

This article details AReaL-Ascend v1.0.5's LoRA RL capabilities, explaining how low-rank adaptation reduces memory overhead for large model reinforcement learning, describing two Megatron LoRA weight update modes (adapter sync vs. merge), and covering cross-node LoRA RL, MoE support, XCCL communication, and Qwen3.6-27B examples for practical deployment.

AReaLAscend NPULoRA
0 likes · 6 min read
AReaL v1.0.5 LoRA RL: Low-Rank Adaptation for Accessible Large Model RL on Ascend
Architects' Tech Alliance
Architects' Tech Alliance
Sep 6, 2026 · Artificial Intelligence

xDeepServe on CloudMatrix384: Full-Stack Design for Large-Scale MoE Model Serving

This article details Huawei's xDeepServe system for deploying massive Mixture-of-Experts models on the CloudMatrix384 supernode, covering the XCCL communication library, FlowServe decentralized serving engine, Transformerless execution architecture with prefill-decode and MoE-Attention decoupling, and hierarchical fault tolerance, achieving 2400 tokens/s per chip at 50ms TPOT.

Ascend 910CCloudMatrix384EPLB
0 likes · 18 min read
xDeepServe on CloudMatrix384: Full-Stack Design for Large-Scale MoE Model Serving