Why the Dawn 8000 Secured WAIC’s ‘Treasure of the Museum’ Award
The Dawn 8000 (Dengfeng) AI supercluster, China’s first fully domestic 100‑k‑card system, overcame long‑standing compute bottlenecks by integrating native super‑intelligent fusion, six self‑developed chips, a lossless RDMA network, high‑density liquid‑cooled cabinets, and a unified scheduling platform, achieving full‑load operation in its first week at WAIC.
At the 2026 Shanghai WAIC AI Conference, the Dawn 8000 (also called Dengfeng) was officially recognized as a “Treasure of the Museum” and showcased as the world’s first operational 100‑k‑card AI supercluster built entirely with domestic components.
Industry bottlenecks that have long limited China’s compute capacity are identified:
Separate super‑computing (FP64) and intelligent‑computing (low‑precision) clusters requiring duplicated infrastructure and data migration, leading to utilization below 50%.
Exponential difficulty in scaling from ten‑thousand‑card to one‑hundred‑thousand‑card systems, with network congestion, hardware‑failure propagation, cooling, storage I/O, and scheduling problems magnifying at scale.
Heavy reliance on imported high‑end chips and high‑speed interconnects, leaving domestic supply chains vulnerable.
Idle compute resources because software stacks for AI‑for‑Science (AI4S), industrial simulation, and large‑model workloads are missing, causing high‑performance hardware to remain under‑utilized.
Solution architecture of Dawn 8000 combines “native super‑intelligent fusion” with a fully domestic, end‑to‑end R&D chain, eliminating the four pain points. The system integrates:
Six self‑developed core chips covering general‑purpose CPUs, AI accelerators, and interconnects, providing a 100% domestic supply chain.
A world‑first high‑density cabinet design that raises compute density by 20×, using orthogonal full‑electric interconnects, immersion phase‑change liquid cooling, and megawatt‑level power per cabinet.
scaleFabric, a native lossless RDMA network supporting up to 114 k cards, sub‑microsecond latency (<1 µs) and 800 Gbps per port, solving large‑scale incast congestion.
A dual cooling system (immersion phase‑change liquid + lake water) achieving PUE as low as 1.04 and enabling heat‑recovery.
ParaStor distributed storage that ranked first globally in both production‑node and 10‑node IO500 benchmarks, scaling linearly with node count.
Gridview 7.0 unified scheduling platform built on MetaStack + K8s, providing a single stack for both scientific FP64 workloads and AI BF16/FP8/INT8 training/inference.
First‑week operational data showed the cluster running at full capacity: daily average of 150 k jobs, peak of over 500 k jobs in a single day, and integration with the national supercomputing internet handling more than 450 k daily jobs.
Application impact spans more than 20 research and industrial tracks, including:
Biology: 80 k cards accelerated full‑process protein‑folding simulations, compressing drug‑discovery cycles from years to weeks.
New materials: 90 k cards performed 3.16 × 10¹³‑atom DFT simulations.
Fluid dynamics: 88 k cards completed a 3.28 × 10¹⁵‑cell turbulent flow simulation.
Quantum chemistry: 80 k cards computed the ground‑state energy of a 152‑spin FeMoco cluster.
Energy & oil‑gas: domesticized seismic imaging algorithms deployed on petroleum exploration scenarios.
Weather: a next‑generation meteorological model built on SwinUNet_n1 delivered high‑precision forecasts.
Ecosystem co‑creation plans were announced:
“Ten‑Hundred‑Thousand‑Card Co‑Creator Incentive” for research teams and tech startups, offering tiered compute subsidies, dedicated scheduling resources, free datasets, and AI‑agent customization.
“Intelligent‑Agent Ecosystem Recruitment” for AI developers, providing a collaborative framework, million‑scale test‑compute environment, and revenue‑sharing for deployed agents.
“Core‑Partner Recruitment” for hardware and software vendors across the full stack, granting exposure, access to a national compute network, and joint development of integrated solutions.
Industry comparison highlights Dawn 8000’s advantages over traditional domestic ten‑thousand‑card clusters and foreign closed‑source hundred‑thousand‑card systems:
Computing architecture: unified FP64‑INT8 full‑precision capability versus split or proprietary ecosystems.
Supply chain: six domestically designed chips provide 100% control, unlike partial import dependence or fully closed foreign stacks.
Network: scaleFabric RDMA offers linear expansion to 114 k cards with sub‑microsecond latency, surpassing congested domestic networks and proprietary IB solutions.
Cooling efficiency: immersion phase‑change liquid cooling achieves PUE 1.04, far better than typical 1.25+ in other solutions.
Storage: ParaStor leads IO500 globally, while competitors suffer from I/O bottlenecks.
Usability: Gridview 7.0 unifies scheduling and provides AI development tools, reducing the high operational overhead of dual‑cluster management.
Accessibility: integrated with the national supercomputing internet, allowing on‑demand nationwide access, unlike isolated private clusters.
These signals demonstrate that China now possesses a fully domestic, large‑scale AI compute platform capable of supporting trillion‑parameter models and Gordon‑Bell‑class simulations, turning high‑end compute from an exclusive research asset into a widely accessible infrastructure.
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
AI Architecture Path
Focused on AI open-source practice, sharing AI news, tools, technologies, learning resources, and GitHub projects.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
