Tagged articles

MoE optimization

2 articles · Page 1 of 1
Huawei Cloud Developer Alliance
Huawei Cloud Developer Alliance
Aug 14, 2026 · Artificial Intelligence

How Huawei Cloud Optimizes the New Xiaohongshu Open‑Source Model dots3‑note Preview for Multimodal Inference

Huawei Cloud adapts and optimizes the 280B‑parameter dots3‑note preview model released by Xiaohongshu, detailing its mixed sparse attention, MoE enhancements, multimodal pre‑training pipelines, unified reward training, and a suite of inference‑time optimizations—including FlashComm, FUSED_MC2, and MTP speculative decoding—delivered via the vLLM Ascend framework for high‑throughput, low‑latency deployment.

Huawei CloudMTP Speculative DecodingMoE optimization
0 likes · 7 min read
How Huawei Cloud Optimizes the New Xiaohongshu Open‑Source Model dots3‑note Preview for Multimodal Inference
Tencent Technical Engineering
Tencent Technical Engineering
Jan 23, 2026 · Artificial Intelligence

Unlocking AI Infra: Distributed Inference, PD Separation, TileLang, and Next‑Gen Agent Infrastructure

This article surveys the 2025 AI infrastructure landscape, covering distributed inference with PD‑separation, dynamic DOPD scheduling, AFD attention‑FFN disaggregation, high‑bandwidth cross‑machine communication libraries, the TileLang programming model, RL train‑inference decoupling via SeamlessFlow, and secure, low‑latency agent infra designs for future large‑scale models.

AI InfrastructureAgent SystemsGPU communication
0 likes · 27 min read
Unlocking AI Infra: Distributed Inference, PD Separation, TileLang, and Next‑Gen Agent Infrastructure