Tagged articles

vLLM Ascend

2 articles · Page 1 of 1
Huawei Cloud Developer Alliance
Huawei Cloud Developer Alliance
Aug 14, 2026 · Artificial Intelligence

How Huawei Cloud Optimizes the New Xiaohongshu Open‑Source Model dots3‑note Preview for Multimodal Inference

Huawei Cloud adapts and optimizes the 280B‑parameter dots3‑note preview model released by Xiaohongshu, detailing its mixed sparse attention, MoE enhancements, multimodal pre‑training pipelines, unified reward training, and a suite of inference‑time optimizations—including FlashComm, FUSED_MC2, and MTP speculative decoding—delivered via the vLLM Ascend framework for high‑throughput, low‑latency deployment.

Huawei CloudMTP Speculative DecodingMoE optimization
0 likes · 7 min read
How Huawei Cloud Optimizes the New Xiaohongshu Open‑Source Model dots3‑note Preview for Multimodal Inference
Huawei Cloud Developer Alliance
Huawei Cloud Developer Alliance
Apr 2, 2026 · Cloud Native

How Kthena Enables Production‑Grade LLM Inference on Kubernetes

This article analyzes the cloud‑native challenges of deploying large‑model inference on Kubernetes and presents Kthena’s architecture—ModelServing, Router, Autoscaler, and ModelBooster—along with Volcano integration, vLLM‑Ascend setup, and a real‑world Qwen3‑235B deployment case, highlighting performance gains and future directions.

Cloud NativeKthenaKubernetes
0 likes · 13 min read
How Kthena Enables Production‑Grade LLM Inference on Kubernetes