Day0 Adaptation of Kimi K3 on Alibaba Cloud Lingjun Zhenwu M890 Supernode
On July 27, Alibaba Cloud announced that its Lingjun Zhenwu M890 supernode instance has been successfully adapted to run the 2.8‑trillion‑parameter Kimi K3 model, achieving a 35 % reduction in first‑token latency, a 1.8× increase in decode throughput, and support for up to 1 M token context through joint chip, software‑stack and framework optimizations.
On the evening of July 27, Alibaba Cloud confirmed that its Lingjun Zhenwu M890 supernode instance is now compatible with the 2.8‑trillion‑parameter Kimi K3 model. The adaptation leverages a joint optimization of the chip, inference platform, and model, resulting in a roughly 35 % reduction in first‑token latency and smoother user experience in long‑context scenarios.
The Kimi K3 model, the latest flagship from Moonshot AI, uses a Mixture‑of‑Experts (MoE) architecture and requires massive compute, memory, and high‑bandwidth interconnect. Traditional AI clusters cannot meet the required high‑throughput, low‑latency, and cost‑effective inference demands. According to Moonshot AI’s official blog, the model should be deployed on a “64‑card‑plus accelerator” supernode to avoid cross‑node bottlenecks.
Lingjun Zhenwu M890 is built on the Pingtouge (T‑Head) training‑inference‑one‑chip solution, featuring the ICN Switch 1.0 interconnect chip that provides 800 GB/s All‑to‑All bandwidth and 9 TB of memory. This enables the entire MoE traffic of a trillion‑parameter model to stay within a single supernode, eliminating inter‑node communication overhead.
Three key optimization areas were addressed:
EP expert‑parallel communication optimization : Using the ICN64 high‑bandwidth domain, the All‑to‑All expert routing was deeply tuned, keeping the 2.8‑trillion‑parameter expert traffic inside one supernode.
MXFP4 mixed‑precision inference : Leveraging the native MXFP4 low‑precision compute of M890, the model’s accuracy remains unchanged while computation density and memory utilization improve significantly.
KDA mixed‑architecture operator optimization : The Kimi Delta Attention (KDA) combines linear and full attention. Operators were aligned with the chip’s parallel compute and memory characteristics, and fused to reduce kernel launch and intermediate memory costs, dramatically boosting attention compute efficiency.
Performance gains observed on the Day0 run include:
First‑token latency (TTFT) reduced by ~35 % compared with the initial baseline.
Decode throughput (tokens per second) increased by ~1.8×.
Mixed‑precision MXFP4 quantization further improves overall inference efficiency.
Thanks to the KDA hybrid architecture, inference compute scales almost linearly with sequence length, reliably supporting up to 1 M tokens of context for long‑document understanding and complex reasoning.
This Day0 adaptation reflects a full‑stack collaboration among Alibaba Cloud, Pingtouge, and the Kimi team, covering chip, inference framework, and cloud platform. The joint effort not only provides an open software stack for trillion‑parameter models but also offers a domestically accessible high‑compute solution. Future work will continue to deepen model adaptation, performance tuning, and engineering deployment to further improve the cost‑performance ratio of Kimi series models on Lingjun supernodes.
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
