Industry Insights 15 min read

Breaking the Scale Barrier: Sugon 8000's 100K-Card AI Supercluster Engineering Battles

The article details how Sugon 8000 overcame engineering challenges to build China's first fully domestic 100,000-card AI supercluster, covering chip architecture, mechanical structure, high-speed networking, storage, and liquid cooling innovations through cross-team collaboration and iterative testing.

Architects' Tech Alliance
Architects' Tech Alliance
Architects' Tech Alliance
Breaking the Scale Barrier: Sugon 8000's 100K-Card AI Supercluster Engineering Battles

1. The Multi-Subsystem Challenge of 100K-Card Scale

Scaling from Sugon 7000's 60,000 cards to Sugon 8000's 100,000 cards is not a simple matter of adding more chips. The expansion introduces "communication walls," "memory walls," and "power walls" — engineering limits where overall performance is determined by multiple tightly coupled subsystems: top-level architecture and chip design, precision mechanical structure, high-speed interconnect network, storage system, and liquid cooling. Any shortfall in one subsystem constrains the entire machine.

Top-level and chip architecture — compute source: Chief designer Li Bin and team anticipated future workload, process, and engineering needs, innovating a "super-intelligence fusion" route where one system supports both traditional scientific simulation and AI large-model training. Chief architect Zhong Yuyi's team reused functional modules within limited chip die area to serve both scientific and intelligent computing tasks, while solving multi-chip high-speed interconnect challenges and filling a domestic gap in switch chip technology.

Precision mechanical structure — hardware carrier: Facing tens of thousands of compute nodes, the structural team controlled component assembly precision to 0.1–0.2 mm. To address deformation risks from multi-layer board stacking, they borrowed the load-bearing logic of traditional Chinese dougong bracket joints, separating load-bearing and connection functions to achieve stable high-density assembly in tight spaces.

High-speed interconnect network — cluster nervous system: The network team built China's largest microsecond-level high-speed network, linking over 10,000 compute nodes. A single link failure would halt large-scale jobs, making the network the communication foundation for 100,000 chips working in concert.

Storage system — data hub: The storage team managed a massive array of 100,000 hard drives handling nearly 100 million read/write requests per second, compressing data latency to sub-millisecond levels. If data access congests, compute chips idle, wasting compute capacity.

Liquid cooling system — hardware operation guarantee: The cooling team tackled the enormous heat dissipation of 100,000 chips at full load. One hour of full-load operation generates heat equivalent to burning 12 tons of standard coal. The liquid cooling system must efficiently remove this heat to sustain continuous full-load operation.

2. Breaking the Scale Trap: Cross-Team Collaboration on Unknown System Problems

The 100K-card cluster is a first-of-its-kind domestic exploration; many faults cannot be predicted by simulation and only emerge when scaling up real tests. Such system-level problems require multi-disciplinary cooperation — forming a closed loop of problem discovery, solution decomposition, and joint testing.

Network: Millisecond-Level Autonomous Failover

During network R&D, small-probability signal transmission errors were amplified at 100K-card scale, with calculations showing ~2.7 transmission errors per hour. Traditional centralized scheduling would take up to 5 minutes for fault correction, causing entire computing tasks to fail. Inspired by autonomous pedestrian traffic at city intersections, the team abandoned fully centralized scheduling and gave network devices autonomous adaptability. Implementation required the chip team to upgrade switch hardware switching capability, the software team to design routing strategies balancing local response and global optimality, and the high-speed signal team to reduce underlying fault probability. After nearly 100 engineers collaborated across disciplines, extreme "manual link disconnection" tests verified the feasibility of millisecond-level autonomous failover.

Storage: "Super Tunnel" Architecture for Billion-Level Independent Channels

Despite previously winning first place on the global IO500 high-performance storage benchmark, the existing shared-channel architecture suffered severe channel contention under 100K-card concurrency. Slow requests dragged down overall response, preventing the massive drive array from delivering its hardware performance. Borrowing the independent-lane design of automated logistics warehouses, the team developed "Super Tunnel" technology, realizing one-to-one dedicated links between CPU, NIC, and each drive. The project simultaneously overhauled storage hardware, drivers, and scheduling software. After multiple stress-test rounds confirming clear paths and performance gains, the solution was validated.

Liquid Cooling: Diamond Copper Direct Die Welding

After selecting diamond-copper composite material for its superior thermal conductivity, full-load system tests still fell short. The team boldly proposed removing the chip's metal lid and directly welding the cold plate to the bare die — a process with no existing reference parameters. Any voids or bubbles at the weld interface would drastically degrade cooling. Materials, thermal design, and process engineers collaborated, repeatedly tuning welding sequence, pressure, and temperature, using X-ray scanning to inspect weld quality. Engineers shuttled samples to fabrication shops overnight, compressing iteration cycles and gradually eliminating defects.

3. Iterative Refinement: Scientific Spirit in Every Debug Cycle

Super-engineering miracles do not happen overnight. The documentary records numerous R&D scenes: system test crashes, cooling missing targets, fatal welding voids, poor high-speed signal quality. Innovation lies in confronting failure and continuous polishing.

High-speed network R&D had near-zero error tolerance; any deviation jeopardized the 100K-chip collaboration goal. The team tested, diagnosed, and fixed in parallel, continuously expanding cluster scale for validation. They observed signal quality via eye diagrams, scanned parameters item by item, and underwent multiple rounds of hardware/software optimization, ultimately achieving millisecond-level automatic failover.

Storage testing frequently faced system overload crashes. The team completely dismantled the old shared-channel architecture, repeatedly rehearsed link logic, and ran tests under real client read/write pressure until performance nearly doubled. The resulting "Super Tunnel" technology, capable of forming billion-level independent high-speed channels, further powered Sugon's centralized FlashNexus and distributed ParaStor storage products to multiple global benchmark tops.

Liquid cooling was a race against time. After selecting diamond-copper from dozens of candidate materials, system cooling still missed targets. Facing delivery countdown, the team challenged the bare-die direct welding process, iterating 10–20 times over 50 days to lock in qualified process parameters — a "buzzer-beater" that completed the final critical piece of the whole machine.

R&D personnel do not fear exposing new problems; every fault and failure becomes the basis for the next optimization. From chip architecture and structural design to network, storage, and cooling, the entire system matured through cycles of debugging, overturning, and re-verification.

4. Specialized Division of Labor, Whole-Machine Collaboration

Sugon 8000's development proves that 100K-card massive compute does not equal simply stacking high-performance hardware. Each subsystem must both excel in its own domain — pushing single-item technical metrics — and mutually adapt and constrain to resolve cross-module contradictions, turning paper specifications into stable, usable real compute power.

Specialized division lets teams focus on core missions: chip team builds the super-intelligence fusion compute source; structural team cracks high-density assembly; network team erects the 100K-card high-speed communication foundation; storage team breaks massive IO concurrency bottlenecks; liquid cooling team achieves extreme thermal management. Yet subsystems inter-constrain: network latency affects storage read/write efficiency; total power rise forces cooling upgrades; cooling hardware changes then constrain whole-machine wiring and structural layout. No matter how excellent a single metric, if it cannot adapt to the whole-machine system, it delivers no real value.

Therefore teams cannot work in silos. After a new architecture is proposed, upstream and downstream must jointly evaluate feasibility; when a subsystem version is ready, it must enter whole-machine joint debugging; when cross-domain faults appear, multi-disciplinary engineers jointly review to locate root causes. Through continuous collaborative polishing, the native super-intelligence fusion chip architecture, dougong-style high-density structure, millisecond-level autonomous failover network, billion-level "Super Tunnel" storage, and diamond-copper direct-weld immersion phase-change liquid cooling all landed, together constituting a fully self-developed supercompute system.

5. Summary and Outlook

The successful development of the Sugon 8000 100K-card AI supercluster is a exemplary practice of large-scale national system engineering. This system addresses current 100K-card compute demand, while its high-speed network and other subsystems lay groundwork for future 200K-card and even million-card clusters. Reviewing the journey — from Sugon-1 breaking foreign technology blockades to Sugon 8000 achieving 100K-card super-intelligence fusion — over thirty years of accumulated technology and talent form the bedrock of China's independent compute breakthrough. Compute competition has evolved into a systemic contest; generations of researchers press forward amidst challenges, solidifying the foundation of China's autonomous compute infrastructure and providing a powerful platform for major scientific discovery and industrial upgrading.

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

storage architecturesystem engineeringliquid coolingsupercomputinghigh-speed interconnectSugon 8000100K-card clusterAI supercluster
Architects' Tech Alliance
Written by

Architects' Tech Alliance

Sharing project experiences, insights into cutting-edge architectures, focusing on cloud computing, microservices, big data, hyper-convergence, storage, data protection, artificial intelligence, industry practices and solutions.

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.