From Usable to Cost‑Effective: How WAIC’s Supernode Expo Showcases China’s Domestic Compute Leap
The 2026 WAIC supernode carnival revealed that Chinese AI compute is shifting from single‑card focus to system‑level solutions, with Huawei, ZTE, Qingwei, Biren, Muxi, New H3C and Sugon demonstrating scalable, cost‑efficient architectures that promise both performance and economic viability.
At the recently concluded 2026 World Artificial Intelligence Conference, the H2 Technology and Innovation Pavilion became the focal point as Huawei unveiled the Atlas 950 SuperPoD prototype, Sugon displayed the 100‑k‑card "Shuguang 8000" AI supercluster, and Qingwei Intelligent highlighted its reconfigurable, switch‑less supernode, turning the once‑conceptual "supernode" into the event’s hottest keyword.
From "Single‑Card Worship" to the "System Era"
Earlier WAIC editions centered on large‑model and AI‑chip parameter races. This year, soaring compute costs, unproven business models, and fragmented deployment scenarios made profitability a more pressing question than model size. The industry is moving from a training‑centric era to an inference‑centric, agent‑application era, where the decisive factor is the efficiency of the whole computing system rather than a single chip’s performance. The China Academy of Information and Communications Technology explicitly states that "supernodes will become the core computing unit of the AI era."
Domestic Supernode Landscape
Huawei : The Atlas 950 SuperPoD uses a 64‑card cabinet as a basic unit, interconnects up to 1,024 NPU cards, provides a unified 256 TB memory address space, and can scale the Atlas 950 SuperCluster to 500,000 cards for trillion‑parameter model training and massive inference concurrency.
ZTE : In partnership with Xizhi, Biren, Muxi, and others, ZTE built a high‑performance Matrix supernode based on the OEX+dOCS architecture. The project entered the SAIL award TOP30 and showcases a benchmark‑level integration of GPU and optical‑interconnect technologies.
Qingwei Intelligent : Leveraging a self‑developed reconfigurable data‑flow architecture, Qingwei connects 4,096 chips via a Mesh topology with point‑to‑point communication, eliminating any dedicated switch chips. The supernode delivers 5 × 10^19 operations per second and cuts interconnect costs by roughly 90 % compared with foreign solutions, saving millions of yuan per cluster. Deployments span Beijing, Xinjiang, Inner Mongolia, Jiangsu, Anhui, and Heilongjiang.
Biren, Muxi, New H3C : Biren’s next‑gen BR20x GPU enables 1,024‑card scale‑up within a single supernode; Muxi launched the "XiJing" S‑series supernode to solidify a full‑stack domestic compute matrix; New H3C presented the UniPoD S80000 series covering scales from 32 to 16,384 cards.
Sugon : China’s first fully domestic 100‑k‑card AI supercluster, "Shuguang 8000," was completed and showcased, marking the transition of AI infrastructure from ten‑thousand‑card to hundred‑thousand‑card deployments.
Economic Pathways for Supernodes
Interconnect cost reduction : In traditional ten‑thousand‑card clusters, switches and NICs account for over 20 % of total system cost. Qingwei’s mesh‑direct, switch‑less design lowers interconnect expenses by about 90 %, translating to multi‑million‑yuan savings per cluster.
Energy‑efficiency gains : Huawei’s Ascend supernodes have been deployed at scale in more than 20 industry scenarios. Qingwei’s reconfigurable data‑flow architecture achieves a transistor utilization rate exceeding 70 %, compared with less than 40 % for conventional designs, meaning more computation per hardware unit.
Ecosystem collaboration : Qingwei is one of the few vendors whose FlagOS stack fully supports core components across non‑GPU architectures, ranking second only to Huawei’s Ascend in non‑GPU scale. It already supports day‑0 adaptation for over 200 large models, including DeepSeek and Qwen, enabling a "write once, run on multiple chips" workflow that dramatically cuts migration costs and development barriers.
Conclusion
The collective supernode surge demonstrates that Chinese domestic compute power has moved from mere availability to economic viability. System‑level innovations—from Huawei’s massive scaling, ZTE’s cross‑vendor integration, Qingwei’s switch‑less architecture, to Sugon’s hundred‑k‑card milestone—answer the core question of delivering AI compute that is not only fast but also affordable and sustainable.
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
Architects' Tech Alliance
Sharing project experiences, insights into cutting-edge architectures, focusing on cloud computing, microservices, big data, hyper-convergence, storage, data protection, artificial intelligence, industry practices and solutions.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
