Day0 Adaptation of Kimi K3 on Alibaba Cloud Lingjun Zhenwu M890 Supernode

On July 27, Alibaba Cloud announced that its Lingjun Zhenwu M890 supernode instance has been successfully adapted to run the 2.8‑trillion‑parameter Kimi K3 model, achieving a 35 % reduction in first‑token latency, a 1.8× increase in decode throughput, and support for up to 1 M token context through joint chip, software‑stack and framework optimizations.

Alibaba Cloud Infrastructure
Alibaba Cloud Infrastructure
Alibaba Cloud Infrastructure
Day0 Adaptation of Kimi K3 on Alibaba Cloud Lingjun Zhenwu M890 Supernode

On the evening of July 27, Alibaba Cloud confirmed that its Lingjun Zhenwu M890 supernode instance is now compatible with the 2.8‑trillion‑parameter Kimi K3 model. The adaptation leverages a joint optimization of the chip, inference platform, and model, resulting in a roughly 35 % reduction in first‑token latency and smoother user experience in long‑context scenarios.

The Kimi K3 model, the latest flagship from Moonshot AI, uses a Mixture‑of‑Experts (MoE) architecture and requires massive compute, memory, and high‑bandwidth interconnect. Traditional AI clusters cannot meet the required high‑throughput, low‑latency, and cost‑effective inference demands. According to Moonshot AI’s official blog, the model should be deployed on a “64‑card‑plus accelerator” supernode to avoid cross‑node bottlenecks.

Lingjun Zhenwu M890 is built on the Pingtouge (T‑Head) training‑inference‑one‑chip solution, featuring the ICN Switch 1.0 interconnect chip that provides 800 GB/s All‑to‑All bandwidth and 9 TB of memory. This enables the entire MoE traffic of a trillion‑parameter model to stay within a single supernode, eliminating inter‑node communication overhead.

Three key optimization areas were addressed:

EP expert‑parallel communication optimization : Using the ICN64 high‑bandwidth domain, the All‑to‑All expert routing was deeply tuned, keeping the 2.8‑trillion‑parameter expert traffic inside one supernode.

MXFP4 mixed‑precision inference : Leveraging the native MXFP4 low‑precision compute of M890, the model’s accuracy remains unchanged while computation density and memory utilization improve significantly.

KDA mixed‑architecture operator optimization : The Kimi Delta Attention (KDA) combines linear and full attention. Operators were aligned with the chip’s parallel compute and memory characteristics, and fused to reduce kernel launch and intermediate memory costs, dramatically boosting attention compute efficiency.

Performance gains observed on the Day0 run include:

First‑token latency (TTFT) reduced by ~35 % compared with the initial baseline.

Decode throughput (tokens per second) increased by ~1.8×.

Mixed‑precision MXFP4 quantization further improves overall inference efficiency.

Thanks to the KDA hybrid architecture, inference compute scales almost linearly with sequence length, reliably supporting up to 1 M tokens of context for long‑document understanding and complex reasoning.

This Day0 adaptation reflects a full‑stack collaboration among Alibaba Cloud, Pingtouge, and the Kimi team, covering chip, inference framework, and cloud platform. The joint effort not only provides an open software stack for trillion‑parameter models but also offers a domestically accessible high‑compute solution. Future work will continue to deepen model adaptation, performance tuning, and engineering deployment to further improve the cost‑performance ratio of Kimi series models on Lingjun supernodes.

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

inference optimizationlarge language modelMoEMXFP4Kimi K3KDALingjun Zhenwu M890
Alibaba Cloud Infrastructure
Written by

Alibaba Cloud Infrastructure

For uninterrupted computing services

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.